Methodology

Every number we publish starts as something someone said

Kompa publishes figures like “80,135 discussions” and “NSR 87.3%”. This page shows how one discussion is counted, what gets discarded before counting, and the formula behind each metric.

This is the shared method behind every report Kompa publishes.

  1. CollectThu thập
  2. ClusterGom cụm
  3. AuthenticateXác thực
  4. UnderstandHiểu nội dung
  5. VerifyKiểm chứng

One batch of data moving through the five processing steps

  1. 01Collect

    Crawling facebook.com

    Millions of records processed daily

    • LinhTikTokXài hai tuần rồi, pin trâu thật@linhchi.ng · 3 hours ago
    • HuyForumGiá này so với bên kia thì ổn áp@huy.tran · 4 hours ago
    • QuyênFacebookNhân viên tư vấn nhiệt tình, sẽ quay lạiFound by meaning

    GlossarySemantic Retrieval

    Found without the brand name in it

  2. 02Cluster

    • NamFacebookGiao hàng nhanh, đóng gói cẩn thận lắm@nam331 · 5 hours ago
    • NamFacebookGiao hàng nhanh, đóng gói cẩn thận ạ 😍96% similar
    • NgọcThreadsĐặt tối qua sáng nay có hàng luôn@ngocanh__ · 6 hours ago

    1 Duplicate

    GlossarySemantic ClusteringNear-duplicate

    A few words apart is still one discussion

  3. 03Authenticate

    • TuấnFacebookApp hay lỗi lúc thanh toán@tuan.ng89 · 7 hours ago
    • shop_giare247FacebookINBOX SHOP GIÁ SỈ 0909 ***@shop_giare247 · 6 hours ago
    • MaiForumTư vấn viên trả lời hơi chậm@mai.phuong · 9 hours ago
    • kèo thơm hôm nayThreadsvào bio nhận ngay 200k@keothom_vip · 8 hours ago

    2 Spam

    Coordinated same pattern · 4 min apart

    GlossaryAuthenticity ScoreCoordinated Posting

    Two accounts, same pattern, four minutes apart

  4. 04Understand

    Original postNhà mạng nào wifi ổn định nhất khu Quận 7 vậy mọi người?

    BảoForum2 hours ago

    Wifidạo nàychập chờnquá,aicũngbịhaymình mìnhvậy ta

    Aspects

    1. Network qualityWifi · chập chờn
    2. Stabilitydạo này · chập chờn · quá
    3. Coverageai · cũng · bị

    Intent

    1. Questionhay · mình mình · vậy ta

    GlossaryAspect-based SentimentUser Intent

    Each aspect gets its own sentiment

  5. 05Verify

    Review threshold70%

    • BảoWifi dạo này chập chờn quá, ai cũng bị hay mình mình vậy ta

      Machine label

    • QuyênNhân viên tư vấn nhiệt tình, sẽ quay lại

      Machine label

    Traces back to source items

    GlossaryConfidenceEscalation Threshold

    Below the confidence threshold, a person reads it

Coverage

Where we listen

360° coverage runs from mainstream media to open data across the internet — social platforms, press, e-commerce, broadcast and forums. Brand conversation happens in all of them, and a number drawn from only one of them is a partial number.

20+Collection platforms

Facebook, TikTok, YouTube, Threads, forums, online news…

  • 200,000+YouTube channels
  • 1,350+Online, print and press titles
  • 200+E-commerce sites
  • 50+TV and radio channels

Public data only

Coverage is half of it. The other half is finding the right discussion inside that coverage: beyond keyword matching, the system retrieves by meaning, so a post that is about the brand without naming it is still in the data.

Plus Facebook pages, groups and community posts across Vietnam. “Platform” and “source” are different units: one platform such as YouTube holds hundreds of thousands of sources.

Process

From raw data to a single number

Five steps, running continuously. Most raw data never becomes a counted discussion — the bulk of it is removed at the duplicate and spam gates.

  1. Identify sources and retrieve

    The system identifies sources relevant to the brand, its products and the configured keyword set. Beyond keyword matching it retrieves by meaning, so a post that is plainly about the brand while naming none of its keywords is still found.

  2. Collect in real time

    Crawlers collect from the identified sources continuously. Each source is recorded with an authority tier: a licensed press title and an anonymous forum account are both collected, and are not equal evidence.

  3. Cluster and de-duplicate

    Near-identical items are grouped into one cluster by meaning rather than by string equality, so a repost with an emoji added, a hashtag appended or a clause reworded is still recognised. One release published across many outlets counts as one discussion with many sources.

  4. Authenticate sources

    Spam, sales solicitation and automated behaviour are removed. Accounts publishing near-identical content inside a narrow window are grouped as coordinated, counted once, and reported separately. Every item carries an authenticity score; anything below threshold is excluded from headline figures and disclosed.

  5. Understand and verify

    Each discussion is segmented, assigned aspects, sentiment per aspect and the writer's intent. Every label carries a confidence score; below threshold it goes to a human reviewer. Every figure in a report resolves back to the item set that produced it.

Collection system

  • Collection across 20+ platforms
  • Automatic classification
  • Millions of records processed daily
  • Real-time updates, 24/7
  • Multi-language support
  • Automated proxy and IP rotation

System

Four processing layers

Data passes through four layers: collection and semantic retrieval, Vietnamese understanding at the level of aspect and intent, verification by confidence score and human review, and finally alerting when something moves.

  1. Collection and retrieval

    The system listens across mainstream news, forums, e-commerce sites and social platforms. Retrieval runs on both the keyword set and on meaning, so relevant discussion is still found when the writer never names the brand.

  2. Vietnamese understanding

    Segmentation, aspect extraction, sentiment per aspect and the writer's intent. The cases specific to Vietnamese — sarcasm, teencode, English code-switching, regional vocabulary — are handled at this layer rather than skipped.

  3. Verification and review

    Every machine label carries a confidence score. Below threshold, the item goes to the 24/7 review team. The quality of each label class is measured against a Vietnamese gold set rather than described with adjectives.

  4. Alerting and management

    When negative discussion passes a threshold, or signs of coordinated posting appear, the system raises an early alert so the team can act before it becomes a crisis.

Humans in the loop

The machine labels. People review.

Sentiment is tagged automatically, then reviewed by people continuously, 24/7. This is why a Kompa report is not the raw output of a model.

Vietnamese is a language where sarcasm, regional slang and abbreviation shift between user groups. Vietnamese-language machine learning handles most of it; the rest needs a reader.

Formulas

Every metric has a formula, not an estimate

Every ratio in a Kompa report comes from a formula stated below. Kompa does not publish a black-box composite score: a proprietary index cannot be audited by the client it is sold to. Drag the slider to see the formula respond.

Net Sentiment Rate (NSR)

NSR64.0%

Neutral discussion is not in the formula — NSR compares positive against negative only. NSR is computed per aspect; the overall figure is the volume-weighted mean of the aspect figures.

Intent mix

  1. Complaint34%
  2. Question26%
  3. Praise18%
  4. Comparison12%
  5. Purchase intent7%
  6. Churn signal3%

The share of discussion by what the writer wanted: complaint, praise, purchase intent, comparison, question, churn signal. Two brands with the same NSR can have entirely different intent mixes.

Sentiment / attribute index

Used to compare a brand’s sentiment or attribute against its category. A value of 1 means level with the category; above 1 means ahead of it.

Positioning

Brand Health Matrix

The matrix places a brand using three social-listening metrics: total mention volume, net sentiment rate (NSR), and audience scale.

Pick an aspect to see its own NSR. A brand can lead on delivery and trail on price — the overall figure is the volume-weighted mean of the aspects, so it hides exactly that.

Worked example
x
Horizontal — total mentionsRelative scale — mention volume is not comparable across categories
y
Vertical — NSR (%)
Bubble size — audience scaleSmaller audienceLarger audience
α — the brand being reported on
β–θ — other brands in the category
NSR 87.3% — the figure this page explains

Hover or tap a zone to read what it means.

The brands plotted here illustrate how to read the matrix. They are not figures from a published study.

Glossary

45 terms used across every report

This is Kompa’s standard vocabulary. When a report says “share of voice” or “unique voice”, it means the definition below and not some other reading of it.

45 terms

Report vocabularyMethod vocabulary

Report vocabulary28

Conversations containing keywords directly or indirectly related to the subject under analysis — across posts, videos, news, comments and shares.

All conversations about the subject under analysis, generated across news media, forums and social platforms.

A group of products, services, business types and organisations sharing a structure, operating toward the same commercial purpose or distributed product, and serving the same customer base.

Words or phrases used to assemble the dataset.

The sources that generate and hold the conversations to be collected and analysed.

Conversations posted and managed by the administrators, plus member comments, on a company’s official website, fanpages and YouTube channels.

Discussions created organically by internet users that mention the brand.

Discussions tagged to a content group or related topic, capturing the character of the information or event.

The users who create posts, or who leave comments.

The relative weight of conversations about a brand, product or service across channels, measured against competitors in the market.

The ratio between positive and negative discussion. Neutral discussion does not enter the calculation.

User emotion on social platforms, in three classes:

  • PositiveTích cựcLiking or loving the brand or campaign.
  • NeutralTrung lậpMentions the brand without expressing emotion.
  • NegativeTiêu cựcDisliking or rejecting the brand or campaign.

Used to compare a brand’s sentiment or attribute against the category as a whole. A value of 1 means level with the category; above 1 means ahead of it.

The emoji reactions users leave on social platforms: like, love, sad, angry, haha.

User behaviour and trends, drawn from the collected data.

The topics carrying the highest conversation volume.

The users generating the most conversation on social platforms.

Discussions or exchanges between two or more people and the interactions they leave: likes, comments, shares.

Conversations about online and offline promotional activity.

The value of the brand across social channels.

Brand health in media, determined by three metrics: total mentions, net sentiment rate (NSR), and unique voice.

The key factors securing the success of a brand’s communications activity.

The process a customer moves through while buying and using a product or service.

Responses and discussion arising directly from a campaign’s message.

Responses and discussion arising directly from the brand and from its campaign messages.

The cost of placing media articles and advertising slots on TV and radio.

The brand value earned from each communications activity across TV, radio slots and media articles.

Method vocabulary17

Finding data by what a sentence means rather than by keyword match alone. A post that is plainly about the brand while naming none of its keywords is still found.

Happens at step 1 · Collect

Grouping discussions that carry the same content into one cluster by meaning, rather than by matching them character for character.

Happens at step 2 · Cluster

Two discussions that are not identical but carry the same content — an emoji added, a hashtag appended, a few words changed. Counted as one discussion.

Happens at step 2 · Cluster

One release or article republished across many outlets counts as one discussion with many sources, rather than as many separate discussions.

Happens at step 2 · Cluster

How strongly an item is believed to come from a real, independent user. Anything below threshold is excluded from headline figures and disclosed in the report.

Happens at step 3 · Authenticate

Several accounts publishing near-identical content inside a narrow window. Grouped, counted once, and reported separately from organic discussion.

Happens at step 3 · Authenticate

Sentiment assigned to each aspect within a discussion rather than to the post as a whole. One post can be positive about delivery and negative about price at once.

What the writer wanted: complaint, praise, purchase intent, comparison, question, or a churn signal. Two brands with identical NSR can differ completely here.

The habit of dropping English into Vietnamese sentences — “ship nhanh”, “sale off”, “oke em”. Code-switched sentences are segmented and labelled like any other.

The abbreviations and respellings common on Vietnamese social media — “ko”, “dc”, “vs”, “iu”. Normalised to their full forms before analysis.

A sentence that means the opposite of what it says. Machine and human readers can disagree here, so it is measured on its own rather than folded into a general error rate.

Happens at step 4 · Understand

Resolving a mention to the thing it refers to: a brand, a product, a competitor or a person. “Shop này”, “nó” and abbreviations all resolve to the right entity.

How certain the system is about a label it assigned, on a 0–100 scale. Carried by every automatic label.

The confidence level below which an item goes to a human reviewer instead of keeping its automatic label.

Happens at step 5 · Verify

A Vietnamese dataset labelled by hand, used as the standard against which system accuracy is measured. Model quality is measured on this set rather than described with adjectives.

Happens at step 5 · Verify

How often two independent human annotators give the same answer. Where people cannot agree with each other, a machine cannot be held to a higher standard.

Happens at step 5 · Verify

Every figure in a report resolves back to the exact set of discussions that produced it. Without this step a number can only be believed, never checked.

Happens at step 5 · Verify

Chart provenance

Every chart carries these six lines

A number with no measurement scope cannot be checked. This is the minimum caption attached to every chart in a Kompa report.

Measurement window
01/01/2026 – 26/03/2026
Platforms counted
Facebook, TikTok, YouTube, forums, online news
Keyword set
12 brand keywords + 34 category keywords
Retrieval method
Keywords + semantic retrieval
Total discussions
14,130
Human-reviewed share
8.4% below confidence threshold, reviewed by hand

Drop any one of the six and the figure stops being comparable — with the previous period, or with another report.

Scope

What this method cannot measure

A method is only useful when its limits are stated. These five apply to every report Kompa publishes.

  • Public data only

    Private messages, closed groups and content deleted before collection are not in the dataset. A discussion that was never public is a discussion that was never counted.

  • Sentiment carries error

    Sarcasm, regional slang and genuinely ambiguous sentences are three cases where machine and human readers can disagree. Continuous review reduces that error; it does not remove it.

  • The keyword set still shapes the number

    Semantic retrieval finds discussion that contains no keyword, but the measurement scope still begins with the keyword set and category chosen. Every report therefore states its window, its platforms and its retrieval method — two reports in the same category with different scopes are not directly comparable.

  • Automatic labels carry confidence, not certainty

    Every machine label carries a confidence score and everything below threshold is reviewed by a person. That narrows the error rather than removing it: what sits above the threshold is still a model estimate, and the accuracy of each label class is measured rather than assumed.

  • Coverage shifts with the platforms

    When a platform changes its API or its visibility rules, coverage in the next period can differ from the last. Material changes are noted in the affected report.

Want this method applied to your category?

Kompa’s analysts will work through the keyword set, the measurement window and the platforms that fit the question you actually need answered.

Talk to the analyst teamBrowse the report library