Methodology
How Public Sentiment works
Public Sentiment measures analysed online discourse. It is not a poll and should not be interpreted as a measurement of voting intention or the opinions of the entire population.
We report “62/100 online sentiment”, “1.2M analysed mentions” or “discussion volume +42%” — never “62% of people support a party”. Mentions are items of content, not people.
Data sources
Content comes only from publicly accessible sources through official APIs or public feeds that permit this use. We do not bypass authentication, rate limits, paywalls or platform protections, and we never collect private messages or private profiles.
- X API v2 — not configured. Set X_BEARER_TOKEN (a tier that includes recent search) and X_QUERIES.
- YouTube Data API — not configured. Set YOUTUBE_API_KEY and YOUTUBE_QUERIES (comma-separated search queries).
- Reddit Data API — not configured. Set REDDIT_CLIENT_ID, REDDIT_CLIENT_SECRET, REDDIT_USER_AGENT and REDDIT_SUBREDDITS.
- Public RSS / news feeds — configured. Set RSS_FEED_URLS to a comma-separated list of https feed URLs you are permitted to use.
- GDELT news monitor — configured. Set GDELT_QUERIES (comma-separated keyword queries; each is limited to Indian sources).
- Instagram (Meta Graph API) — not configured. Requires a Meta app with approved Instagram Graph API / Meta Content Library access. Not implemented until access is granted.
- Facebook (Meta Content Library) — not configured. Requires approved Meta Content Library / Graph API access for public Pages. Not implemented until access is granted.
- Demo data generator (synthetic) — not configured (development only; synthetic data about fictional parties, always labelled DEMO DATA). Set ENABLE_DEMO_PROVIDER=true for development. Never enable in production.
Coverage is therefore uneven: platforms without approved API access (for example Instagram and Facebook until access is granted) are not included, and each platform’s user base differs from the general population.
Collection
Each cycle fetches new items since the last successful fetch for every enabled source, then normalises them (text, timestamps, engagement counts). We store only what analysis needs: the public text, a link to the original, publication time and engagement counts. Account identifiers are replaced by a salted one-way hash used only to count distinct accounts and detect repetitive posting.
Update cycle
The processing target is every 6 hours: collect → normalise → deduplicate → detect language → spam signals → AI classification → aggregation → snapshot. Each source is refreshed at most once per cycle, as its API access, quotas and terms allow. Every page shows when the data was last successfully updated, and warns when the latest update is old. If one source fails, the others continue.
Deduplication
We detect exact copies (content hash), normalised copies (ignoring case, links, punctuation, retweet prefixes), near-duplicates (64-bit SimHash over words, Hamming distance ≤ 6, for texts of 10+ words) and reposts. Duplicates are kept for platform-level volume counts but excluded from sentiment, so one viral post copied thousands of times cannot distort the result. Identical content is classified once and the result reused, which also controls AI cost.
Spam & automated accounts
Automated-account detection is probabilistic; no method can reliably identify bots from public content alone. We record signals — unusually high posting frequency, bursts within a minute, the same account repeating content, identical content from many accounts, abnormal engagement, link-only posts and hashtag stuffing — and combine them into a spam score. Items above 0.6 are excluded from sentiment; lower scores reduce an item’s weight.
AI classification
Each item is classified server-side by an OpenAI model using structured (JSON-schema) output that is re-validated before storage. The model and prompt version (currently political-analysis-v2) are stored with every classification for auditability. The AI labels individual items; it never scores parties — all numbers are computed by our own code. Mentions are matched to parties, people and issues only when unambiguous; uncertain matches are dropped rather than guessed. AI summaries of events and trends are generated only from the collected items, and fall back to purely data-derived text when the evidence is insufficient.
Sentiment
Sentiment is the emotional tone of the text: positive, negative, neutral, mixed or unclear. For parties, people and issues we use the sentiment expressed toward that entity.
Stance
Stance is the author’s position toward a political entity: supportive, critical, neutral, questioning or uncertain. It is classified separately from sentiment: “I hate how people attack Party A” is negative in tone but supportive of Party A; heavy sarcasm (“Brilliant priorities!”) can use positive words with a critical stance.
Topics
Items are mapped to a configurable registry of issues (Economy, Employment, Inflation, Education, …) and also receive short free-form topic labels, so new topics can be discovered without code changes.
Events
An event is created when at least 15 analysed items within 12 hours refer to the same development and discussion of the involved parties/issues runs at least 100% above its prior 24-hour rate. Event pages compare online sentiment before and after the event.
Scoring
Online sentiment = score = 50 + 50 × Σ(weight × value) / Σ(weight) over the last 24 hours, with values positive = +1, neutral/mixed = 0, negative = −1 (unclear excluded). Weights: recency (half-life); engagement (logarithmic, capped at 2×); per-account dampening (1/√posts by same account); spam signals (reduce weight; excluded above threshold); duplicates (excluded). Engagement can never dominate: its effect is logarithmic and capped. A score needs at least 5 unique items, otherwise we show “Insufficient data”. Discussion change compares mentions with the previous 24 hours; discussion growth compares the last 360 minutes with the 360 before. Percentage changes are hidden when the comparison base is too small.
Confidence
Each score carries High / Moderate / Low confidence based on sample size, the share of distinct accounts, the number of platforms, duplicate rate, spam concentration, spread over time and consistency between platforms. The reasons are shown next to the label.
Privacy
We do not collect private messages, passwords, private profiles or unnecessary personal information, and we do not store account handles. Representative posts link back to the original public post, where the platform controls its visibility.
Neutrality
Public Sentiment does not endorse or oppose any party or candidate, predict elections, or infer voting intention. Parties are never ranked as winners or losers; comparison tables can be ordered alphabetically, by discussion volume or by recent change. Party colours are assigned by a fixed neutral palette, not party branding.
Limitations
- Online discourse is not representative of the population; people who post about politics differ from those who don’t.
- Coverage depends on which APIs are configured and their quotas; some platforms are not included at all.
- AI classification makes mistakes, especially with sarcasm, code-mixed language (e.g. Hinglish), regional languages and context-dependent references.
- Bot and coordinated-activity detection is probabilistic and incomplete.
- Small samples are volatile; always check the confidence indicator and sample size.
- Engagement counts are reported by platforms and may change after collection.