Two real findings from this system’s own automation. Both were found, both were fixed, both are in the commit history that produced this page. This is the shape of what a review returns.
The bug. A stylesheet was stored in an ordinary Python string. It contained a CSS escape for the section symbol:
CSS = """
h2::before { content:"\00a7 " counter(sec) ". "; }
"""
Python reads \00 inside a non-raw string as an
octal escape, not as the start of a CSS one. The literal did not
contain the four characters \00a7. It contained a NUL byte
followed by 0a7.
The consequence. The build wrote a NUL byte into the middle of the published HTML. The build reported success. The publication-safety filter — which reads every page before it ships — passed it, because that filter screens for meaning and holds no opinion about whether the bytes are still a web page. A corrupt page was one command from being committed and served, and nothing in the pipeline would have said a word.
The evidence. The output file read as bytes rather than as text. Reading it as text is what hid the problem: terminals, editors and browsers all render a NUL as nothing at all, so the page looked correct in every tool anyone would naturally reach for.
The fix, and the more important fix. One character — the string became a raw literal:
CSS = r"""
That repairs this instance. It does not repair the class, so a permanent guard went into the same check that had waved it through:
bad = sorted({hex(b) for b in page.encode("utf-8")
if b < 9 or 13 < b < 32})
if bad:
sys.exit(f"CORRUPT OUTPUT in {where}: control bytes {bad}")
No control byte reaches the public site again, whatever future escape invents one. That distinction — fix the instance, then close the class — is what the review is for. A review that returns only the first is a spellcheck.
The bug. Seven jobs run unattended on a schedule. Each begins by syncing from version control, behind a deliberate guard: refuse to run rather than act on stale state.
git pull --ff-only --quiet || { echo "refusing: pull failed"; exit 1; }
The guard is correct. The scheduler is also configured to catch up on runs it missed. So when the machine woke from sleep, all seven fired in the same second — into a network stack that was still coming up. Every sync failed. Every job took the guard and exited. Seven for seven.
The consequence. The system did no work for roughly forty-four hours. No alert fired, because every job had exited through a branch working exactly as designed: a failed sync is supposed to stop the run. The watchdog did not fire either — it is a scheduled job, so it was in the same pile, failing the same way, for the same reason. The outage was found because a human looked at the output and thought it seemed quiet. That is not a monitoring strategy.
The evidence. The system journal: all seven units starting in the same second, each reporting an unreachable network, each exiting 1 — and then a clean run on the very next scheduled tick, once the network was up. Two facts together prove it. Simultaneity rules out seven independent faults; self-recovery without intervention rules out a broken configuration. Either fact alone points somewhere less useful.
The fix, and the more important fix. A precondition now runs before each job: wait for the dependency to actually answer, then hand control to the job’s own guard, which stays the arbiter of a genuine outage. That closes the race. It does not close the real finding, which is worse, and which the review states plainly: a fleet of jobs whose failure mode is silence has no monitoring, only the appearance of it. Every job exited nonzero, and nonzero went nowhere. A watchdog that shares a failure domain with the thing it watches is the same bug wearing a hat.
Neither finding came from a test suite. Both systems had passing checks throughout. That is the category of bug this review exists to find.
Free scope check. Send a repository link or describe your workflow in a paragraph. You get back which tier fits, what the review would look at, and an honest answer if the answer is that you do not need it.
Get a free scope check →