Method
Written by: Fatih
The method has three layers and they are deliberately kept apart: observation, archive, reading. What the machine does and what a human does sit on the same page but under separate headings; you should be able to see where one ends and the other begins.
1. Observation — the machine
A monitoring engine pulls a fixed list of sources at regular intervals. Today that list, visible channel by channel in the archive, holds 25 channels spread across seven source families: three open Hugging Face endpoints (trending models, trending datasets, daily papers), the release lists of ten GitHub repositories, three arXiv categories (cs.AI, cs.CL, cs.LG), the release feeds of four PyPI packages, three institutional blog feeds (OpenAI, Google AI, Google DeepMind), Anthropic's announcement sitemap, and the U.S. SEC's EDGAR full-text search.
How much each channel takes is fixed too, and does not move to taste: the top 20 for trending models, the top 10 for datasets and papers, the first 10 for arXiv and blog feeds, the latest 5 on PyPI, at most 5 on GitHub — the limits that shape this are published in GitHub's own rate-limit documentation — the newest 10 Anthropic announcements, and the first 15 filings of the last seven days on EDGAR. The number of records you see on a channel page sits inside that limit.
The engine's rules are narrow, and deliberately so:
- Raw page content is not fingerprinted. Dynamic session tokens would make every run look “changed”; so only meaningful fields are taken (title, link, date, name).
- Links come only from the source's own response. Deriving an address from an identifier counts as fabrication; if the source gives no address, the record carries no link. That is why records in some channels have no link — it is the rule, not a gap.
- Dates are never invented. If the source supplies none, only our “first seen” stamp is written; and if we do not know that either, we say so plainly.
- Every record has a stable key. The same record arrives with the same key on the next run, so “new” and “seen again” never get confused.
- Noise is filtered by measurement. A repository that ships dozens of builds a
day (
llama.cpp) collapses to a single row; one that stamps many package tags per release (vercel/ai) drops to one row per package. The filter deletes nothing; it limits how many rows that run may write. - If even a single channel fails, the whole run is abandoned and that day's state is not written. A half-finished day is a day whose gaps are invisible; rather than publish a page that looks complete on partial data, we publish nothing.
Every channel page spells out all four links of the chain: the source's own page, the endpoint the engine actually calls, that endpoint's official documentation, and the source's terms of use. The endpoint address printed on the page is compared against the address in the engine's code on every check; if the two diverge, the gate turns red and publishing stops. So “we take this from that endpoint” is not a promise — it is a tested claim.
Why these channels
Three tests were applied. First, the source must publish its own output: a model's release list, an organisation's announcement feed, a regulator's filing archive. Second-hand news digests are not on the list, because when they are wrong we cannot show where the error came from. Second, the endpoint must be publicly open — you can call every address shown here yourself and get the same answer; a claim hiding behind a closed data vendor cannot be audited. Third, the channel must carry dates: a signal without a date cannot answer the archive's central question, “when did we know?”
Those tests deliberately leave popular signals out. Social-media mention counts, closed-formula “attention scores” and token prices are not on the list — the first two because they cannot be audited, the third because we do not want to say anything about price itself.
2. Archive — append only, never delete
When a record is seen for the first time it is appended to the archive as a single line and is never rewritten. That way the question “when did we know” cannot be edited after the fact. The archive is open to everyone; no subscription required.
First seen is the moment we first saw a record; it need not match the source's publication date, and the archive shows the two separately.
This rule has a measured cost, and we do not hide it: because archive files are never rewritten, records that existed before monitoring began were inherited without a first-seen stamp, and that gap is permanent. Filling it in later would mean writing a date we do not actually know. Next to those rows you will see “first-seen date unknown — inherited from an earlier baseline”.
3. Reading — the human
The person who reads the observations, connects them and writes what they might mean is named. A reading is an interpretation and is labelled as one. You can walk back to every observation an interpretation rests on.
Every judgement published in the newsletter carries at least one source address, and this is not a habit left to good intentions: the number of source-less judgements is counted before publication, and if it is not zero the issue does not go out. You may disagree with a reading; you will not have to disagree without seeing what it rests on.
What is measured before publishing
Pages are not uploaded by hand; every publication passes a check, and if the check is red there is no publication. Among the things measured: endpoint addresses matching the engine, the number of source links on channel pages, the value sentence on the home page carrying real numbers, the price being readable inside the raw HTML, all four elements of the publishing calendar being in place, and the author page carrying a direct e-mail address. The check itself is tested too: we separately show that the gate turns red when a rule is broken on purpose, because a gate that never turns red is decoration, not a gate.
Numbers are not hand-written either. The channel and record counts on the home page are computed from live data as the page is generated; written by hand they would go stale unnoticed.
When we are wrong
We do not delete a line. We strike it through, write why, and record the fix on the Corrections page. If you have spotted an error, write to us: if we accept it, it appears on the corrections page; if we do not, we write down why as well.
What we do not do
We do not forecast prices, we give no trading advice, and we do not answer personal investment questions. A record appearing here carries no positive or negative judgement about that project or company. All of these limits and the reasons behind them are on the Limits page.
Source notices
The notices our sources ask for stand here and on the relevant channel page as visible text — not merely embedded in the page source, because as we measured machine readers ignore embedded data. Our monitoring requests carry an identifying user agent, so who we are shows up in your server logs; if you own a source and see a problem with how we use it, every option including removing the source from the list is on the table.