Service Updates

NewSumm Update: Honest Source Counts, Fewer Duplicates and Better Categories

PN
Paige Newsom
7 min read
2 views

We audited how NewSumm groups and categorises the news against a week of real data. Source counts now count outlets rather than feeds, duplicate stories are merged automatically, sections are chosen by reading the story, and long-running stories are told in linked chapters.

NewSumm reads around 20,000 articles a day from over a thousand news feeds, groups the ones about the same event into a single story and summarises what each outlet reported. The grouping is the whole product: if it is wrong, readers see the same story three times, or a story that claims many sources but really has one. This week we audited how well it works against a week of real production data, and shipped the fixes. This post covers what we found, what changed, and what you will notice.

The short version:

  • Source counts are honest. One outlet's several feeds, or one wire story reprinted word for word, now count as one source.
  • Far fewer duplicates. Stories about the same event are grouped more reliably, and duplicates that slip through are merged automatically.
  • Better categories. A story's section is now chosen by the AI that writes its summary, not by keyword matching.
  • Long-running stories are told in chapters. Each chapter has an up-to-date summary and links to the coverage before it.

What the audit found

We took a copy of a week of production data and checked each stage of the pipeline by hand and with an offline replay of the clustering. Four problems stood out.

Many "multi-source" stories had one source

A story is only shown once at least two sources cover it. But a "source" meant a feed, and many outlets publish several feeds. One large UK tabloid alone had four, and some outlets had a dozen. Groups of papers under one owner also publish the same wire copy word for word. Nearly half of the stories we showed as covered by several sources were really one outlet, or one article reprinted.

The same event appeared as several stories

Articles were grouped one at a time, in the order they arrived. When the first two reports of an event were worded differently enough, each started its own story, and the two grew side by side without ever being joined. Around one visible story in six had a near-identical twin elsewhere in the feed.

Comparing article pairs by hand showed that the threshold for "same story" was set too strictly: pairs well below it were almost always the same event. Raising recall that way on its own would only have made the duplicates problem worse, so it had to come with a way to merge stories after the fact.

About a quarter of stories were in the wrong section

Categories were assigned by matching keywords, and the matching looked for words inside other words. "Who" filed stories under health, "tory" matched story and history, and "ashes" matched crashes. A sample we checked by hand had a tennis final under politics and a flood under health. The category was also fixed by the first article in a story and never revisited.

Some stories never ended

Every new article on a topic could join an existing story, however old. Recurring items such as daily puzzle hints and lottery results grew into stories months long, and each new article pushed them back to the top of the feed. Long-running news had the same problem: one election story covered two months of coverage under a single summary.

What changed

Sources are counted by outlet

Every feed now belongs to a publisher, worked out from where its articles actually live rather than from its feed address, since many feeds are served through third-party feed hosts. A story's source count is the number of distinct publishers, capped by the number of distinct headlines, so the same article syndicated to three papers counts once. Headlines are compared on their letters and digits only, so the different ways feeds encode an apostrophe do not split them apart.

When we applied this to the existing week of stories, about 40% of the stories shown as multi-source dropped out of the feed. Every one we checked was a single outlet or a reprint.

Better grouping, with automatic merging

The similarity threshold is now set from the hand-checked comparisons rather than guessed. After each clustering run, stories whose overall content has converged are merged: the smaller story's articles, bookmarks and reading history move into the larger one, and its old address redirects permanently to the story it was merged into, so shared links and search results keep working.

Categories chosen by reading the story

The AI that writes each summary already reads the whole story, so it now also chooses the section, from a fixed list with a short definition of each. The category is updated every time the summary is, so it follows the story as it develops. Keyword matching still gives a story a provisional section before its first summary, and now only matches whole words.

Stories have a lifespan, and chapters

A story now accepts new articles for five days after it is first seen. Coverage after that starts a new story with its own summary, and the two are linked. A story page lists the earlier coverage it continues, and an older story points readers to the latest one, which matters because readers arriving from search often land on older pages.

Summaries are also now written from a story's newest coverage, with each outlet getting a place before any outlet gets a second. Previously they could be written from the first articles to arrive, so a story that ran for days could still be describing how it first broke.

Growing stories no longer disappear

One bug turned up along the way. When a new article joined a story, the story was marked as needing a new summary in a way that also removed it from the feed until that summary was written, which could take up to half an hour. The busiest stories were the ones most often missing. A story now keeps showing its current summary until the new one replaces it.

The results

We tested the changes by running the full pipeline overnight on a local copy of production, about 10,400 newly fetched articles, and compared the result with the same measures before the changes:

  • Stories shown as multi-source that were one outlet or a reprint: 46% before, 0.1% after.
  • Genuinely multi-source stories per 10,000 articles: about 390 before, about 850 after.
  • Visible stories with a near-duplicate: 18% before, about 3% after.
  • Stories in the wrong section (hand-checked samples): about 27% before, about 5% after.
  • Stories still growing after several days: 27 of the top 200 before, none after.

The category figures come from small samples checked by hand, so read them as a direction rather than a precise rate. The duplicates that remain are mostly an older chapter next to its newer one for a few hours, until the older one moves down the feed.

What you will notice

  • Fewer stories in the feed for a short while after the change, since the ones that were really a single outlet are gone. They are replaced over the following hours by genuinely multi-source stories.
  • Big stories cover more outlets, and the same event appears once.
  • Sections that match what stories are about.
  • On long-running stories, an "Earlier coverage" list on the story page, and a link to the latest chapter on older ones. The same is coming to the NewSumm app in its next update.

If you spot a story in the wrong place or a duplicate that should have been merged, we would like to hear about it.

Share this article

PN

Paige Newsom

Author at IfHighLow

Related Articles

Figaro's Big Proud Balloon

New Story: Figaro's Big Proud Balloon

Once upon a time on a sunny island, lived Figaro the Frigatebird. He had sleek black feathers, huge pointed wings, and a big bright red throat pouch. This pouch could puff up like a balloon! But, Figa...

NN
Nora Newsworthy
Koa's Curious Day

New Story: Koa's Curious Day

Koa the Kea awoke to a bright morning in the mountains. His olive-green feathers sparkled under the sun, and he stretched his wings to reveal flashes of orange. With a curious glint in his bright eyes...

CC
Clara Chronicle
Gita's Gentle Gift

New Story: Gita's Gentle Gift

In a peaceful river, Gita the Gharial glided gracefully between the reeds. Her long, slender snout was perfect for catching fish. The sun sparkled on her olive-green scales, and her golden eyes glimme...

PH
Preston Herald

Stay Updated

Get the latest insights, tutorials, and industry news delivered to your inbox.

We respect your privacy. Unsubscribe at any time.