HEPHFORGE

Sources

Written by: Fatih

Every source we watch is public and none of them is paid for. What we take, and just as importantly what we deliberately do not take, is written out here — the blind spots of an intelligence product matter as much as its field of view. This page is not a statement of intent; it describes what is written in the engine's code. Every number below was read from that code or from the archive itself.

What we watch today

As these lines are written the engine watches 25 channels, and the public archive holds 36 channel files and 1,621 record lines, of which 190 sit inside the live window. The full list of channels, what each one takes and what it does not take, is written out channel by channel in that same archive.

A channel page carries four addresses at once: the source's human-readable page, the endpoint the engine actually calls, that endpoint's official documentation, and the source's terms of use. These are not decoration. Our gate compares the endpoint address printed on the page against the address in the engine's code on every run, and refuses to publish if the two have drifted apart.

How much we take per channel

The intake ceilings are written into the code and cap how many records a single round may pull. On the Hugging Face side we ask for 20 records for the model list and 10 each for datasets and daily papers; the endpoints themselves are described in the Hub API Endpoints documentation. On the GitHub side we pull at most 10 releases from each of ten repositories and write at most 5 of them to the channel after the noise filter; the limits that shape this are published in GitHub's own rate-limit documentation.

On the arXiv side we take 10 items per round for three categories (cs.AI, cs.CL, cs.LG) and print the citation text the source asks for wherever the source is used; the conditions are set out in the arXiv API Terms of Use. On the PyPI side we read the release feeds of four packages, and the feed format is defined in PyPI's own feed documentation. On the SEC side the query is fixed and the window slides: 8-K filings from the last 7 days are read through the EDGAR full-text search interface.

The source sets our pace, not us

How many requests per minute a source accepts is stated in its own documentation, and that number — not our preference — sets the tempo of a round. This is why we run one round a day. Running more often is within our power; staying inside the source's limit is the price of taking data from it.

We ask while saying who we are

Every request we make carries an identifier that says who we are: HEPHFORGE and a reachable email address. That is the name under which we appear in the source's server logs, and a source that is unhappy can write to us directly. A crawler that hides itself cannot be accountable to the source it watches. On the SEC side this is not courtesy but a requirement: an unidentified request is refused — the rule is published by the SEC itself — and our own probe tests that behaviour on every run.

The limit we keep on news feeds

From news feeds we take only the title, the link and the date. Neither summary nor full text is stored or republished; a reader who wants the piece follows the link to the publisher's own page. This is both the legally correct and the product-honest choice: we are not republishing the news, we are recording when something appeared.

We do not invent addresses

A record's address comes only from the source's own response. Hugging Face's response carries no address field, so records from those channels have an empty address — deriving an address from an identifier would not be reading data but manufacturing it. An empty field is more honest than a filled-in guess.

The noise filter, and its formula

A source that publishes very often will fill a whole channel by itself unless filtered. We run three rules: the ggml-org/llama.cpp repository, which produces dozens of builds a day, contributes only its single newest record per round; the vercel/ai repository, which ships many package versions in one release, keeps one record per package; and the remaining repositories are capped at 5. The filter deletes nothing. It caps the lines written that round, and what it caps is stated both here and on the channel page. An exclusion whose formula cannot be shown produces something as unquestionable as the record itself.

A round happens completely or not at all

If even a single channel errors during a round, that day's state is never written: the generator, the gate and the publishing step are all three skipped. The price is a visible gap. The alternative — publishing whatever happens to be at hand and rendering the gap invisible — turns into a quiet lie in a publication that talks about measurement.

This behaviour was tested by reality on the morning of 5 September 2026: the EDGAR endpoint returned a server error, the engine stopped, and that day's round was abandoned — which is why no record line for that date exists in the archive. The following day the round ran normally. We did not backfill the gap; because we cannot, it stands written.

Dropped records stay with us

Most sources return a “most recent N” window; once an item falls out of it, the source no longer shows it. It remains visible here, because the job of an archive is to remember what the source forgot. Of the archive's 1,621 lines today, 131 have no first-seen date: these are records inherited from before watching began, and their dates are not invented after the fact — they stay empty.

What we deliberately do not use

A closed channel is not deleted

We narrowed our field once: the channels belonging to the crypto period were shut down. The shutdown record, that period's engine and that period's state were not deleted; they were frozen and kept. The answer to “what were we looking at, and when” should not be a thing that can be edited afterwards.

How a new channel gets added

A source being interesting is not enough to make it a channel. Four things are required, in order: its terms must be readable and permissive, the endpoint to be called must have official documentation, the response must contain a stable identity field that lets a record be recognised across rounds, and the channel's own probe must be written. No channel opens before all four are done.

The third item eliminates the most candidates. In a source without a stable identity, a “new record” and an “old record with a changed title” blur into one another and the archive inflates on its own every round. Such a channel produces a pile that looks right but cannot be counted — and a pile that cannot be counted is of no use here.

Every number on this page is measured again

The numbers here are not written once and left. They are counted again from the archive before each publication, and if they have changed they are written as they now stand. A page that keeps carrying an old number is the quietest kind of wrong.