How stenograf works
stenograf is an open-source-intelligence (OSINT) monitor for public US news. It crawls a curated list of outlets, enriches the content, and presents it as a feed, a map, cross-source events, entity dossiers, and coverage comparisons.
Sources
All underlying material comes from real, named US outlets, and every article links to its original. The full list of monitored sources and their health is on the Sources page.
What is computed algorithmically
These stages are deterministic and use no language model:
- Named-entity extraction (people, organisations, places) from the US text.
- Folding the spellings and grammatical cases of one person or organisation into a single entity. This is automatic and can over- or under-merge, so a mention count is a best-effort figure rather than an exact one.
- Linking entities to Wikidata: where an entity page carries a one-line description, it comes from Wikidata — not from our coverage, and not from a language model.
- Geocoding against a curated gazetteer (map pins come from a vetted list of places).
- Clustering articles into events and scoring popularity by the number of independent sources.
- Topic and sub-topic classification from the article's own text. Where the classifier is not confident it leaves the topic blank — we would rather abstain than guess.
- Extractive US summaries and full-text search.
- Stance: how far one outlet's telling of a story sits from how the others told it.
- Signal detection: a statistical comparison of a recent window against a rolling baseline. Severity is banded from that score — no human and no language model judges it.
What a language model generates
This content is produced by a large language model and is marked with an AI analysis badge:
- Cross-source synthesis (what sources agree on and where they diverge).
- Coverage comparison — US vs World — divergence scores, omissions, and tone.
- Entity dossiers for people and organisations.
What you see without an account
Without signing in, the browsing surfaces — feed, search, and the event and entity lists — show only the last 30 days of coverage. Individual story and entity pages stay readable at any age, so a bookmark or a search-engine result keeps working. If an older story seems missing from a listing, this is usually why.
An important caveat
Divergence scores, lists of omissions, and tone/framing labels are automated interpretation, not editorial fact. Language models can be wrong, biased, or incomplete. Treat generated content as a starting point and always check the linked original articles.