How KhabarCheck works
KhabarCheck (खबर चेक, “news, checked”) is a fully automated news aggregator. There is no newsroom: software fetches stories from the sources below every half hour, groups articles that describe the same event, and computes a reporting-confidence score for each story. We never write news — every headline links to the outlet that published it.
Automated doesn’t mean unaccountable. The feed is generated without a human in the loop, but the system isn’t: reader reports, methodology changes, and anything that looks like a high-risk error are reviewed by the person who built and maintains it — publicly, since the repository and every reader-filed issue are open for anyone to read. Resolved corrections are listed on the corrections log.
The reporting-confidence score (0–100)
It measures the strength and diversity of available reporting — not whether every claim in the story is true. Seven outlets can repeat the same wrong claim; one excellent reporter can break a true story alone. Read the score accordingly.
| Component | Weight | What it measures |
|---|---|---|
| Corroboration | 40 | How many independent reporting origins a story has — not a raw outlet count. Outlets under shared ownership (Mint + Hindustan Times), copies of the same wire report (PTI, ANI, Reuters), and articles that merely cite another outlet’s reporting collapse into one origin. One origin scores low; five score full marks. Known limit: we analyse headlines and summaries, so several outlets independently covering the same single statement can still overcount — flagging single-statement stories is planned. |
| Source reliability | 30 | The average rating of the outlets involved, from a hand-maintained per-outlet score kept in the source code. |
| Primary source | 15 | Whether the story links to an official source — government releases, court records, regulators, or international bodies. |
| Headline check | 15 | An automated check for clickbait, sensationalism, and opinion presented as news (an AI model when available, keyword heuristics otherwise). |
Bands: Widely reported is 75+, Partially corroborated is 50–74, Limited reporting is below 50. A story covered by a single lower-rated outlet is marked Not independently corroborated regardless of its number.
The score estimates how well-corroborated a story is right now — it is not a truth verdict. Breaking news often starts as Not independently corroborated and climbs as more outlets confirm it. Stories that surfaced within the last two hours from a single outlet carry a Developing tag: they appear early by design, and their score should be expected to move.
How the feed is ordered
The home feed is ranked, not purely chronological: a story’s position comes from its recency (half-life of ~18 hours), how many outlets corroborate it, and a category weight — tech and sports are down-weighted (×0.6) so a gadget launch never outranks a court verdict. By default the ordering carries no built-in preference for India or World stories either way; use those tabs to focus on one. Nothing is ever removed by ranking: every ingested story keeps its page, appears in region and topic filters, and ships in the sitemap and RSS feed.
Automated is not the same as neutral. Recency half-life, the tech/sports discount, corroboration weight — each is a choice about what matters more, made by one person, encoded in software instead of an editor’s desk. No ranking of news is neutral in an absolute sense; ours doesn’t claim to be. What we can offer instead is transparency: every weight above is published here and in the source code, specifically so it can be read, argued with, and disputed — rather than trusted on faith.
Ongoing story hubs
A single running story — a protest, a court case, a crisis — can produce a dozen separate developments in a few days: health updates, political reactions, celebrity comments, court orders. Each is its own verified story with its own score, but showing all of them as separate feed entries reads as repetition. When four or more stories share a specific, non-generic detail — not just a topic like “Kerala High Court,” which rules on unrelated cases daily — they collapse into one ongoing story card showing the latest development, linking to a page with the full timeline. Nothing is hidden: every development keeps its own page, score, and place in the region/topic filters and sitemap; the hub only changes how the home feed displays them together.
Known limits: grouping runs on shared names and places extracted from headlines, not on understanding what a development actually says — it can occasionally miss a development or, rarely, group two similar-sounding but unrelated stories. Splitting developments into background / reactions / claims-checked (as opposed to one flat timeline) needs deeper analysis of each article and isn’t built yet.
Sustainability & our promises
KhabarCheck currently runs on a few dollars a month and carries no ads and no accounts. As it grows, keeping it running may mean adding revenue — clearly labeled sponsorships or ads, reader support, or optional accounts for features like saved topics. We would rather be honest about that possibility than make promises we might have to break. One promise is permanent: we will never track our readers — no behavioural profiling, no selling data, no third-party trackers, whatever the revenue model. Anything optional (like an account) will stay optional, and everything commercial will be labeled as such.
Media Blindspots
Every outlet is assigned an editorial bucket: Indian mainstream (large commercial outlets), Indian independent (non-profit or independent newsrooms), international, and official (government sources). A story is flagged as a mainstream blindspot when it is India-relevant, has zero Indian-mainstream coverage, and is reported by at least two non-mainstream outlets (or one plus community discussion). A blindspot is a coverage gap, not an accusation — stories can be early, niche, or simply missed.
Check a forward
The check tool matches pasted text against our last 7 days of stories, entirely on your device — nothing you paste is sent anywhere or stored. A match shows who is reporting the claim and how strongly it is corroborated; no match means unconfirmed by the outlets we track, not necessarily false.
Under the Radar
Stories with strong engagement on Reddit or Hacker News but little mainstream coverage. This surfaces news the big outlets have not picked up — sometimes because it is early, niche, or inconvenient; sometimes because it is wrong. Community posts are treated purely as discovery signals and never count toward corroboration.
Source reliability table
Ratings (0–100) are opinions, seeded from public press-reliability research and maintained in the open — the full list ships with the site’s source code so anyone can audit or dispute it.
Honest limitations
- Wire services (AP, Reuters) no longer publish free feeds; their reporting reaches us indirectly through outlets that syndicate them.
- Twitter/X is not included — its API pricing is beyond this project.
- Story grouping is automated and occasionally merges or splits stories incorrectly.
- India vs. World tagging relies on keyword matching and can occasionally mistag a story — most often a niche foreign business/tech story that names a company but not a country.
- Non-news service content — horoscopes, puzzle answers, lottery results, multi-event digests — is dropped at ingestion; it isn’t reporting.
- The reliability table is a maintained opinion, not an objective fact.