What AstroPress should refuse to publish
A publishing engine needs enforceable rules for what reaches readers.
Wolfe Services · · 5 min read
The easy demonstration of a content engine is a new article appearing on a screen. Give it a topic, retrieve some material, generate a draft, and the system looks productive.
The more useful demonstration is a draft that fails to publish for a specific reason. Perhaps its supporting source is missing. Perhaps a product capability has been described as a measured result. Perhaps the article was edited after the review was recorded.
That is the direction we are taking with AstroPress and the editorial workflow behind these field notes. A publishing platform needs an explicit account of what is allowed to reach readers. Generating more words increases the value of that account.
Begin with a question we have something to say about
A content calendar can easily become a list of familiar marketing topics. Law firms do ask about SEO returns, website ownership, intake speed, and attribution. Those are reasonable subjects. Familiarity with the question does not supply an original answer.
Our topic briefs therefore require an angle and a distinction from existing Wolfe coverage. Before drafting, the engine searches the editorial corpus for related articles and evidence. If the site already makes the argument, the next piece needs a different purpose or the existing article needs an update.
The corpus is intentionally selected. It contains Wolfe articles and curated summaries of operational evidence. It does not need private client mail to explain a measurement failure. Selecting sources is an editorial decision before it is a retrieval problem.
The useful question for the drafter is: what can we explain here because we built, investigated, or measured something? If the answer is only “this keyword seems valuable,” the brief is not ready.
Retrieval supplies candidates for inspection
A vector search can find a passage whose wording differs from the query but whose subject is related. That is helpful in a body of work where the same system may be described as an intake bridge, a lead handoff, or an attribution boundary.
It can also retrieve a fluent statement that is outdated, incomplete, or simply wrong. A similarity score measures a relationship between representations of text. It is not a confidence score for the truth of a claim.
Our source ledger keeps a dated summary, its limitations, and references to the inspected originals. The original remains the place to check the claim. A generated article that cites another generated article has not gained independent corroboration merely by adding a link.
The Wolfe workflow also records a version for its indexed corpus. If the approved material changes, retrieval requires the index to be refreshed. That addresses a defined technical failure mode: drafting against an old representation of the available evidence. It cannot establish that the evidence itself is correct.
Editing needs its own pass
We use a separate humanizer pass after drafting. Its job is to remove formulaic constructions, vague claims, repetitive section openings, and the voice of a generic marketing summary. It should make the reasoning easier to follow.
That pass has strict limits. It cannot add a plausible anecdote to make the article feel personal. It cannot remove a meaningful caveat because uncertainty sounds less confident. It cannot turn an internal observation into a claim that someone independently audited the result.
For these notes, the useful voice is an operator explaining a decision: what happened, why it mattered, and what remains unresolved. An editor should be able to point to the source of the experience. Personality added by fabrication would defeat the purpose of publishing the work.
After automation produces an editable version, the review has to examine its argument and evidence. An article can be grammatically clean and materially misleading.
Put the publication state in the system
Astro’s content collections support schemas for the metadata attached to content. We use that capability to keep publication state explicit. A generated draft is stored outside the public collection while it is being worked on. A draft imported for review remains excluded from article routes, topic listings, related links, and RSS.
A published field note needs a review record tied to its final content and metadata. If either changes, the record becomes stale and the build check rejects it until reviewed again. That is narrower than a claim that software can prove someone checked a fact. The record is an editorial attestation; the hash checks whether it still refers to this version.
A practical publication contract looks like this:
| Stage | Required decision or evidence |
|---|---|
| Topic | A reader question and a distinct angle |
| Retrieval | Related coverage and selected evidence |
| Draft | A developed argument with visible limits |
| Edit | A specific prose review that preserves the facts |
| Publication | Explicit status and a review of the final version |
| Later revision | A new review when the text materially changes |
A failed build can protect the reader
The same idea already appears in the Law Tracker work. Its generator stops when a digest produces too few parsed items, rather than quietly publishing an empty ledger. That threshold does not verify every legislative interpretation. It catches a recognizable failure in the data supply.
Good publishing controls are similarly specific. They reject known bad states and make unresolved questions visible to the editor. Google’s guidance on helpful content also asks publishers to consider original value and the basis for trust; a production quota alone does not answer those questions.
If you are evaluating a content engine, ask for a demonstration of a draft with a missing source, a stale review, and an unpublished status. Find out what happens to each. Then ask who examines a sentence that passes every mechanical check but still overstates the evidence. The answer to that last question is part of the publishing system too.