When a news data provider tells you they monitor 800,000 sources, that number sounds significant.
The question worth asking is: 800,000 sources as of when, and how many of them delivered anything useful last week?
TL;DR
- A source count tells you what a provider has configured, not what’s actively delivering content.
- Traditional coverage benchmarks reward weak crawling — more duplicates, dead sources and noise inflate the numbers.
- The metrics worth asking for: headline diversity, within-domain diversity, and live output by time window.
- If a provider can’t tell you how many sources were active last week, that’s the answer.
A source list is a configuration snapshot. It may include sites that went dark two years ago, domains split across dozens of subdomains, feeds counted separately from their parent sites, and publications that technically exist but produce no content worth indexing. None of that is necessarily dishonest.
Some of it is simply poor housekeeping. The effect on your evaluation is the same either way.
Coverage benchmarks in the news API market have a structural problem: the easiest way to improve them is also the wrong way. A crawler that collects duplicate URLs, keeps dead sources on the books, and indexes page furniture as article text will score well on almost every traditional benchmark.
More sources. More articles. More search matches.
Better numbers, worse output.
This piece explains why that happens, what to measure instead, and what transparent coverage reporting actually looks like.
The trap: better crawling can make the numbers look worse
Consider what happens when you fix a duplicate URL problem. You stop collecting the print version, the mobile version and the AMP version of the same article as three separate items. Your article count drops. Your source count may drop too, if some of those were registered as separate feeds.
By traditional benchmarking logic, you’ve made your coverage worse.
In practice, you’ve made it more useful. The corpus is smaller and less repetitive. Downstream deduplication is cheaper. Searches return fewer false positives.
This is the trap.
The metrics that are easiest to report are also the most sensitive to the things a good crawler actively removes. If you’re evaluating providers on raw volume alone, you may be selecting those with the weakest extraction and deduplication.
Seven ways coverage numbers get inflated
These aren’t always deliberate. Weak deduplication, weak source hygiene and weak extraction each produce the same effect on the numbers.
A useful exercise when reviewing a provider’s coverage claims is to ask which of these could be contributing.
- One site becomes many: sections, subdomains, regional editions and separate feeds from the same publisher counted as distinct sources.
- The zombie source list: dead or inactive sites kept on the books because removing them would reduce the headline number.
- Same article, several URLs: tracking parameters, print versions, mobile versions and AMP pages treated as separate articles.
- Old news becomes new: bad date parsing or update timestamps misread as publication dates, making archive content appear freshly published.
- Non-news becomes news: topic pages, gallery pages, listing pages and SEO landing pages indexed alongside editorial content.
- Synthetic network publications: publications where the URL works after the hostname is replaced, indicating auto-generated content farms rather than independent editorial sources.
- Page furniture becomes article text: menus, sidebars, related story links and tag clouds indexed as article body text, inflating match counts.
A memorable example: one “missing source” that appeared on a provider’s published source list was only findable on Google because Google had indexed the leaked list itself. The source had never been crawled.
The better benchmark: diversity metrics
Volume asks, “How much?” Diversity asks, “How much is different?”
Collecting more copies of the same content increases volume.
Finding more distinct content increases diversity.
They are not the same thing, and only one of them is what you’re actually paying for.
Three diversity metrics are particularly useful for evaluating a news API:
Within-domain diversity
The ratio of distinct domain-plus-headline combinations to total articles. This captures a specific problem: volume that’s inflated by repetition within the same publisher rather than by repetition across publishers.
Repetition across publishers is sometimes valuable; the same story picked up by multiple outlets is a meaningful signal. Repetition within a single publisher is much more suspicious. A high within-domain diversity score indicates the crawler isn’t padding its numbers by collecting the same article multiple times from the same site.
Headline diversity
The ratio of distinct headlines to total articles in a given period. Simple to calculate, hard to inflate. If a crawler is collecting duplicates at scale, this ratio will be low. If the corpus is clean, it will be high.
As a concrete example: twenty million articles in a week sounds impressive. Twelve million distinct headlines say considerably more about what you’re actually getting.
Language and country diversity
The number of languages and countries represented in the actual output, not in the source configuration. A provider might list 200 countries in their source base. The relevant question is how many countries’ sources actually delivered content within your query window.
| Period | Articles | Distinct domain + headline | Within-domain diversity |
|---|---|---|---|
| Hour | 121,305 | 119,447 | 98.5% |
| Day | 2,911,345 | 2,845,948 | 97.8% |
| Week | 20,380,058 | 19,799,202 | 97.1% |
Based on Opoint feed statistics, October 2026. Within-domain diversity = distinct domain + headline combinations / articles.
| Period | Articles | Distinct headlines | Headline diversity |
|---|---|---|---|
| Hour | 121,305 | 86,579 | 71.4% |
| Day | 2,911,345 | 1,809,088 | 62.1% |
| Week | 20,380,058 | 12,154,023 | 59.6% |
Based on Opoint feed statistics, October 2026. Headline diversity = distinct headlines / articles.
In a typical week, more than 97% of Opoint’s articles carry a headline not already seen on the same domain. A crawler that collects print, mobile and AMP versions of the same story would score lower here, not higher. That’s what makes it hard to fake.
Repetition across publishers is different.
The same story picked up by several outlets is a useful signal, so we don’t count it as padding.
Twenty million articles in a week sounds impressive. Twelve million distinct headlines says more about what you’re getting.
This figure falls as the window widens because stories get republished across outlets, so compare providers over the same period. It also depends on whether a provider keeps syndicated and wire copy. Opoint does, and groups it with the equalgroup field. A provider that strips syndicated copies will score higher without crawling any better. For crawler hygiene, read it alongside within-domain diversity.
What live output reporting looks like
There’s a useful distinction between inventory claims and live output.
“We monitor X sources” is an inventory claim, whereas “Here’s what arrived in the feed this week” is live output.
The difference matters because it’s verifiable. You can’t easily audit a source list, but you can audit a feed.
The table below shows what live output looks like when it’s reported by time window rather than as a static total.
| TIME WINDOW | SOURCES | LANGUAGES | COUNTRIES |
|---|---|---|---|
| Every minute | 1,050 | 53 | 91 |
| Every hour | 22,500 | 89 | 170 |
| Every day | 97,000 | 120 | 193 |
| Every week | 145,000 | 130 | 195 |
Based on feed statistics, September 2026. Daily figures vary by news cycle.
Sources are counted only if at least one article was published in the period. A source that published nothing that week doesn’t appear in the weekly count. That’s a meaningful distinction from a source list that counts every site ever configured, regardless of current activity.
The full live coverage data, updated from feed statistics, is available on our coverage page.
The continuous curation question
Live output numbers are only as reliable as the process behind them. A feed that shows 97,000 active sources per day implicitly claims that the curation keeps that number meaningful.
The relevant questions are:
- How quickly are new sources discovered and onboarded?
- What happens when a source changes its publishing patterns or goes inactive?
- How are low-quality, spam and duplicate sources identified and removed?
- Is cleanup periodic or continuous?
These questions don’t have a single right answer, but the answers reveal how stable and trustworthy the coverage numbers are over time. A source count that was accurate six months ago may no longer be accurate today if there’s no ongoing maintenance process to keep it honest.
Three questions to ask any news API provider
When you’re evaluating a news data provider and they present coverage numbers, these three questions cut through most of the ambiguity:
Is it alive?
Did the sources in your count actually deliver content in the relevant period? Can you show me live output data, not just a source list?
Is it different?
What’s your headline diversity score? What’s your within-domain diversity? How many distinct languages and countries appeared in actual output last week?
Is it really there?
When your system reports a match, was the matching text actually inside the article body? Or could it have come from a menu, sidebar, tag cloud or related links section?
These questions reward providers who have invested in clean extraction and deduplication.
They’re harder to answer well if volume has been inflated by weak crawling. That’s exactly why they’re worth asking.
What this means for your evaluation
Coverage benchmarks aren’t going away.
Source counts and article volumes are easy to communicate and compare, which is why providers lead with them. The useful shift is to treat them as a starting point rather than a conclusion.
A high source count with low headline diversity is a warning sign. A large article volume that drops sharply when you filter for distinct headlines warrants investigation. A provider who can’t tell you how many sources were active in the past seven days may not have that figure to hand.
The providers worth working with are those for whom these questions aren’t uncomfortable. If they’re reporting live output by time window, tracking diversity metrics and maintaining the source base continuously, the volume numbers tend to be defensible as well.
Good coverage should be hard to fake. The metrics that make it hard to fake are the ones to prioritise in your evaluation.
Frequently Asked Questions
What is a news API coverage benchmark?
A coverage benchmark is a metric used to compare how broadly a news data API monitors the web. Common benchmarks include total source count, article volume, language coverage and country coverage. These figures vary widely between providers and are not always directly comparable.
Why doesn't source count reliably measure coverage quality?
Source count reflects how many sites a provider has configured to crawl, not how many are currently active. It can be inflated by counting subdomains separately, retaining inactive sources, or splitting feeds from the same publisher.
A provider with 800,000 configured sources may have fewer distinct, active sources than one reporting 200,000.
What is headline diversity in a news API?
Headline diversity is the ratio of distinct headlines to total articles in a given period.
A high ratio indicates clean deduplication; a low ratio suggests the corpus contains many copies of the same content indexed under different URLs.
What questions should I ask when evaluating a news API?
Ask whether the sources in the count actually delivered content recently (live output, not just a source list), what the headline diversity and within-domain diversity scores are, and whether search matches are drawn from article body text or from page furniture such as menus and sidebars.
Want to see how adverse media coverage fits into your compliance workflow across the markets and languages that matter to you?
Want to see the numbers behind this?
We run live coverage reporting by time window; sources, languages and countries are counted only when content actually arrives. If you're evaluating news data providers and want to see what that looks like in practice, we're happy to walk you through it.