The SEO changes we kept, reversed, and couldn’t judge
An experiment ledger needs room for a failed idea and an uncertain result.
Wolfe Services · · 5 min read
An experiment log is easy to maintain while the changes are being shipped. The difficult entry is the verdict, especially when the work took time and the result gives you little to celebrate.
Our JTNY experiment ledger contains all of those outcomes. A title reversion was worth keeping. An attempt to deepen employment content did not achieve the intended movement. An exit-popup change produced too little evidence to separate the alternatives confidently.
A firm paying for ongoing technical SEO should see those verdicts. They explain how the team will use the next month of work. A list of completed edits cannot answer that question on its own.
A decision rule belongs before the result
For a change to teach us much, we need to record the problem it is supposed to solve, the affected pages, the period to read, and the condition that would count as success. Otherwise the report can drift toward whichever number improved.
A page can receive more impressions while attracting no additional inquiries. It can gain clicks from queries outside the intended subject. A changed title can coincide with a new answer block, a search update, or a shift in demand. Choosing a favorable metric afterward makes the intervention look easier to interpret than it was.
The operating discipline is modest: state the hypothesis, preserve the comparison, and decide what evidence would cause us to keep, reverse, or hold the change. Then allow the verdict to be disappointing.
That does not require a claim that every SEO change is a scientific experiment. Most of this work is an observational before-and-after read. Calling it that helps keep a practical decision from growing into a causal claim we cannot support.
The title we put back
One entry concerned a New York court-attire page whose title had been moved toward a jury-duty angle. The later intervention reverted the title. The page-level read improved, and the ledger recorded a win.
That supports the operating decision to keep the reversion. It does not isolate the effect of the title: the page also received an answer block and FAQ during the surrounding work. The verdict explicitly records that confound.
The distinction is useful when explaining the result to a firm owner. We can say that the revised page performed better in the observed comparison and that we chose to retain the change. We cannot honestly turn that into a rule that this title pattern produces a particular lift on every legal page.
Reversing a change can also be good work. An ongoing engagement should give the team room to restore something that was already serving readers well. If every update has to be defended as progress, weak changes acquire a long life.
The extra depth that did not move the target
Another intervention added employment-law depth to a local page. The hypothesis was that the page could move its employment query family into a stronger search position.
The subsequent read did not show the intended movement. A query that appeared strong after the change had already been strong before it. The ledger marked the intervention as lost against its stated objective.
This is where additional word count can mislead an editorial team. A more complete page may be useful for readers without producing the search change someone predicted. Those are separate claims. We should preserve useful content for a stated reason, rather than retroactively treating completeness as evidence that the ranking hypothesis succeeded.
The next decision might be to revisit intent, internal navigation, or the role of the page. It should not automatically be another layer of prose on the assumption that the first layer almost worked.
The popup we could not call
The exit-popup read involved very few events. The destination changed at the same time a timer was removed. Even where a literal success condition could be described as met, the recorded verdict remained inconclusive.
That is the appropriate place to acknowledge the limits of the observation. A small number of clicks and submissions cannot reliably tell us whether one treatment will keep outperforming another. Changing both timing and destination makes the interpretation harder again.
Here is how those reads differ:
| Intervention | Operating verdict | Limit on the claim |
|---|---|---|
| Restore the court-attire title | Keep the reversion | Other page changes prevent isolating title impact |
| Add employment depth | Target not achieved | The strong query already performed before the edit |
| Change the exit popup | Inconclusive | Sparse events and simultaneous treatment changes |
These are summaries of internal reads closed in September 2026. They are not a benchmark dataset or a controlled comparison between websites. We have deliberately omitted headline lift percentages that would hide the limits in a footnote.
Sometimes the denominator changes underneath you
The same operational record documents a separate session-counting defect fixed in September. Before the correction, the site could assign a new identity and session to each pageview. Comparing lead rates per session across that boundary would compare two different counting systems.
That can produce an attractive-looking improvement without establishing that the site persuaded more people to inquire. The correct response is to document the definition change and establish a clean baseline. The traffic investigation explains why we now ask more questions of the denominator.
Ask your team to bring one kept change, one failed hypothesis, and one inconclusive read to the next review. For each, ask what was decided before the work and what will happen next. A provider who can explain those decisions is giving you something more useful than a monthly inventory of edits.