<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/">
  <channel>
    <title>Spoolbound devlog</title>
    <link>https://spoolbound.com/devlog</link>
    <description>Notes from building Spoolbound: the systems, the measurements, and what they changed.</description>
    <language>en-GB</language>
    <lastBuildDate>Sun, 06 Sep 2026 09:00:00 GMT</lastBuildDate>
    <atom:link href="https://spoolbound.com/feed.xml" rel="self" type="application/rss+xml" />
    <item>
      <title>The garage became a bay on a station, and the simulation never noticed</title>
      <link>https://spoolbound.com/devlog/the-garage-became-a-bay</link>
      <guid isPermaLink="true">https://spoolbound.com/devlog/the-garage-became-a-bay</guid>
      <pubDate>Sun, 06 Sep 2026 09:00:00 GMT</pubDate>
      <description>We moved the whole game somewhere else. The renderer needed a reskin, the economy needed nothing at all, and the only thing that really broke was the website — which spent four days describing a place the game no longer had.</description>
      <content:encoded><![CDATA[<p>The game used to be set in a garage. You dug it out, filled it with printers, and grew it into a warehouse. As of a few days ago it is a fabrication bay on a station, the walls are bulkheads, and there is a starfield through the porthole.</p>
<p>That sounds like the kind of change that costs a rewrite. It did not, and the reason it did not is the only genuinely interesting thing about it.</p>
<h2>What actually had to change</h2>
<p>A station module interior is a room with walls and machines on a floor. A garage is a room with walls and machines on a floor. The renderer draws rooms with walls and machines on floors.</p>
<p>So the largest line in the estimate went from <em>rebuild the world</em> to <em>reskin the room</em>, and most of what shipped was surface: wall material, floor material, a porthole with stars behind it, and the strings that name the place. The camera arithmetic was not touched. The decision that one room fills the screen at a time was made months ago for reasons about legibility on a phone, and "one wide deck" became "one wide module" without a single number moving.</p>
<p>The economy did not change at all. Not adjusted — untouched. Printers still make parts, parts still become products, products still fill contracts. A part that warps in open air still warps, and for the same reason it did before, because that reason is about the material and not about the postcode.</p>
<h2>Why it was cheap, stated as the thing it actually was</h2>
<p>There is a rule this project has followed since very early on, for reasons that had nothing to do with any of this:</p>
<p><strong>The simulation imports nothing from the renderer. It uses whole numbers for money and time, and it never reads a clock — the current time is always handed to it.</strong></p>
<p>That was done for testability, and later for a server that does not exist yet. It has been mildly annoying to maintain. Every so often something obvious would be easy if the simulation could just ask what time it was, and the answer was always no.</p>
<p>The payoff arrived in a currency nobody was saving up in. Because the simulation had never been told where it was, it could not object to being moved. There was no code that believed in a garage. The fiction lived in the renderer, in the strings, and in the art, and all three of those are things you can repaint.</p>
<p><strong>A boundary you maintain for one reason pays out for a different one.</strong> That is not a plan you can make. You cannot predict which wall will turn out to be the load-bearing one. What you can do is keep the walls honest, so that when you do want to move the house, you find out it was never nailed to the ground.</p>
<h2>What we got wrong</h2>
<p>The setting changed on the third of September. The website went on describing a garage until today.</p>
<p>Every page said the old thing. The one-paragraph description that search engines and answer engines quote said it. The press kit said it, and offered four screenshots of a place that no longer exists — which is worse than saying it, because a journalist would have used them in good faith and published a picture of the wrong game.</p>
<p>Nothing caught it, because nothing was watching. This repository has a gate that scans the built site for words we have promised not to publish and claims that would become retractions, and it is a good gate: it has caught real problems, including one that had been live for days. But it checks the copy against a list of rules. It has no idea what the game is currently about, so it cannot tell that a true sentence has quietly become a false one.</p>
<p>That is a familiar shape here. We have written before about checks that pass while looking at the wrong thing, and about a number that was measured correctly and was still false because the world had moved underneath it. This is the same species, applied to prose:</p>
<p><strong>A sentence that was true when it was written does not report that it has expired. Somebody has to go and look.</strong></p>
<p>The honest fix is not a cleverer gate. Nothing automatic is going to read a design decision and work out which paragraph on a marketing site it just invalidated. The fix is that changing the premise has a step in it that says <em>and now go and change what you told people</em>, and that step did not exist. It does now.</p>
<h2>What is here and what is not</h2>
<p>The bay, the station it sits in, and the ship growing in the yard are all real and all in the build. The first loop closes in about six minutes from a cold start with nothing granted: clear some floor by hand, take an order for six brackets, wait, collect, deliver, and the wallet reads $75.60. That is measured from a real save, on a phone, with the developer spawn switched off.</p>
<p>There is more in there than this post describes. Some of it is not finished and some of it we have not decided how to talk about, and this devlog's standing position on that is the boring one: anything we have not described has not been announced. Treat the absence of a thing here as <em>not announced</em>, not as <em>not in the game</em>.</p>
<p>The screenshot on the front page is cropped to the floor for exactly that reason. The full frame has interface in it that belongs to a conversation we have not had yet, and showing less seemed better than implying more.</p>]]></content:encoded>
    </item>
    <item>
      <title>Our playtest robot bought machines for ten hours and never switched any of them on</title>
      <link>https://spoolbound.com/devlog/it-bought-them-and-never-switched-them-on</link>
      <guid isPermaLink="true">https://spoolbound.com/devlog/it-bought-them-and-never-switched-them-on</guid>
      <pubDate>Sun, 30 Aug 2026 09:00:00 GMT</pubDate>
      <description>The harness worked, the runs completed, the numbers looked plausible, and a whole column of results turned out to be counting something other than what it was labelled. Here is what survived the correction, and why.</description>
      <content:encoded><![CDATA[<p>The last few of these have been about checks that pass while looking at the wrong thing. This one is a different animal, and worse in a specific way: nothing was checking anything. A measuring instrument was quietly measuring something else, and it produced numbers that looked exactly like the numbers it was supposed to produce.</p>
<h2>What it did</h2>
<p>We have a probe that plays the game as a flawless player — buys the right thing at the right moment, never misses a collection, runs for hours of simulated time in seconds. Everything we know about pacing comes from it, and previous entries here have leaned on it heavily.</p>
<p>Buying a machine and <em>starting</em> it are two separate commands.</p>
<p>The probe sent the first one. It never sent the second.</p>
<p>So every machine it ever bought sat idle for the rest of the run. Ten hours of simulated time, a shop full of equipment, and <strong>not one part was ever printed by any of it.</strong></p>
<h2>Why nothing caught it</h2>
<p>The runs completed. No errors, no crashes, no assertion failed anywhere — the probe deliberately has no assertions, because it is an instrument rather than a test, and I have defended that decision in this devlog before. I still think it is right. But it does mean the only signal an instrument gives you is its output, and the output looked fine.</p>
<p>It looked fine because <strong>the columns were plausible.</strong> A table headed with machine names, filled with non-zero numbers that rose as the run went on. It was counting <em>purchases</em>. Reading it as production required no misunderstanding at all — the label said one thing, the arithmetic underneath said another, and a plausible number is far more dangerous than a missing one. A blank column gets investigated. A column of reasonable-looking figures gets quoted.</p>
<p>And it got quoted. Here, among other places.</p>
<h2>What survived, and this is the part worth keeping</h2>
<p>When the correction landed, most of the shop-level figures moved. But one headline number did not: the estimate of how little a returning player finds waiting for them after a long absence.</p>
<p>It survived because <strong>it had never come from the probe.</strong> It was arithmetic done directly on the catalogue — print duration, buffer size, multiply out — a completely separate method that happened to answer the same question.</p>
<p>That is the whole lesson, and it is cheaper than it sounds:</p>
<p>&gt; <strong>A number you can reach two different ways is a number you can trust. A number with one source is a number you are hoping about.</strong></p>
<p>We did not do that deliberately. We got lucky, in that one figure had been derived twice by accident, and it was the only one still standing when the instrument turned out to be wrong.</p>
<h2>Three more wrong numbers, three different species</h2>
<p>Correcting the probe turned up others in the same sweep, and they are worth separating because the fixes are not the same.</p>
<p><strong>A true number about a thing that no longer exists.</strong> One figure had a research track taking over thirteen hours. It was accurate — for a version of the game that was retired weeks ago, whose progression system no longer exists. Nobody mis-measured anything; the world moved and the number stayed. The current figure is closer to an hour and a half. <strong>Correctly measured, and still false</strong>, which is a category most people do not check for.</p>
<p><strong>A units slip.</strong> One machine's price was out by a factor of a hundred, in the direction that turns a modest purchase into an absurd one. Cents where dollars were meant. It had been sitting in a table being read as a real figure.</p>
<p><strong>A number that was simply better than reported.</strong> The engaged player was recorded as running at around two-thirds efficiency. Re-measured, they are at a hundred percent — the buffer they were thought to be overflowing peaks at roughly half its capacity. That correction went the pleasant direction, and it is worth saying out loud that the sweep found some of those too. A correction pass that only ever finds bad news is a pass that is confirming something.</p>
<h2>The rule this one leaves behind</h2>
<p>The instrument had no assertions and that was the correct design. It reports; a human decides. But an instrument still needs one thing a test does not:</p>
<p>&gt; <strong>Something has to check that the instrument is exercising what it claims to exercise.</strong></p>
<p>Not the values — the <em>behaviour</em>. In this case: did any machine ever transition out of idle? A single check for that would have failed on the very first run, ten hours of nothing would never have been recorded, and several months of quoted figures would never have needed a correction.</p>
<p>That check now exists. It is one line, and it is embarrassing how cheap it was.</p>]]></content:encoded>
    </item>
    <item>
      <title>The build said fine. It shipped a game with no game in it.</title>
      <link>https://spoolbound.com/devlog/the-build-that-said-fine</link>
      <guid isPermaLink="true">https://spoolbound.com/devlog/the-build-that-said-fine</guid>
      <pubDate>Thu, 27 Aug 2026 09:00:00 GMT</pubDate>
      <description>Our release build completed successfully and produced an app with the entire economy missing. Then the check we wrote to catch that skipped, in exactly the state it existed to catch. Both were found the same way.</description>
      <content:encoded><![CDATA[<p>Three posts ago I wrote about checks that return an honest green about something adjacent to the thing you cared about. I did not expect the series to continue this well.</p>
<h2>The build</h2>
<p>Spoolbound's simulation — every price, every contract, every machine — lives in a native library compiled separately from the game and loaded at runtime. That is a normal arrangement and it works.</p>
<p>Producing a release build of the app ran to completion. Exit code zero. One warning in the log.</p>
<p>The app it produced <strong>did not contain the library.</strong> No economy at all: no contracts, no printers, no money. A shop of zeroes.</p>
<p>The warning said so, in the way warnings do, in the middle of a few hundred lines nobody reads when the thing succeeded. The exit code said everything was fine, and the exit code is what anything automated would have believed.</p>
<h2>Why it stayed invisible</h2>
<p>Every mechanism we had was pointed slightly to one side of it.</p>
<p><strong>The debug build was fine.</strong> Everyone develops against debug. It loaded the library correctly, so every day of work confirmed a pipeline that was only half-working.</p>
<p><strong>The release path had never been exercised.</strong> There was no reason to build one until there was a reason to ship one, and by then it had been quietly broken for as long as it had existed.</p>
<p><strong>And success was reported.</strong> This is the part worth sitting with: nothing failed. There was no red anywhere to investigate. A pipeline that returns zero and emits an artefact is a pipeline that every reasonable observer — human or automated — reads as working.</p>
<p>The failure was not that a build broke. It was that <strong>a build could be completely broken and still say fine</strong>, and nothing in the system was positioned to notice the difference.</p>
<h2>The fix, and the rule it needed</h2>
<p>The fix itself is small — a release target, a declaration pointing at the compiled library, and the wiring that registers it.</p>
<p>One detail is worth extracting, because it is a trap rather than a fix. Declaring the library <strong>without</strong> having built it is worse than the state we were in: it converts a loud warning into a silent, empty ship. So the declaration and the binary have to land in the same change, always. A half-applied fix here is a downgrade.</p>
<p>Then the actual protection: a check that runs a <strong>real release build</strong> and asserts the library is present in the finished bundle, by name and by plausible size.</p>
<p>It deliberately trusts <strong>neither the exit code nor the absence of warnings</strong>, because both were measured lying about this exact bug. The only thing it believes is the artefact.</p>
<h2>And then the check did it too</h2>
<p>Mutation-test the gate: remove the library, confirm it screams, restore.</p>
<p>It did not scream. It <strong>skipped</strong>, and reported a cheerful green.</p>
<p>The first version of the check skipped whenever the release library was absent — reasonably enough, on the theory that a machine with no build toolchain should not fail somebody's test run. Except <em>the library being absent is the entire state the check exists to catch.</em> It was written to detect exactly one condition and it had been taught to look away from it.</p>
<p>That is the same bug as the original, one level up, written by someone who had just spent hours thinking about the original. Which is the strongest argument for mutation-testing that I know: reasoning about a check cannot find this. Only breaking the thing it protects and watching what the check does.</p>
<p>It now skips only where there is no toolchain at all. Toolchain present and library missing is a failure. Verified by removing the library and watching it go red, then restoring it and watching it go green.</p>
<h2>The pattern, four times now</h2>
<ul>
<li>An exit code read from the last command in a pipe, not the one being tested.</li>
<li>A <code>noindex</code> flag that filtered a sitemap and emitted no meta, so it controlled nothing.</li>
<li>A test asserting a delete button was greyed out, passing on code that would delete an account anyway.</li>
<li>A build reporting success while producing an empty artefact — and the check for it skipping on the one state it was written to detect.</li>
</ul>
<p>None of these was a wrong answer. Every one was a <strong>correct answer to a question nobody had asked</strong>, and each looked exactly like working software right up until somebody broke something on purpose.</p>
<p>The rule we keep arriving at, from a different direction each time:</p>
<p>&gt; <strong>Before you trust a green, break the thing it protects and confirm it screams.</strong></p>
<p>We have written that down three times now and it has caught something new on each of them. I have stopped treating that as a coincidence.</p>]]></content:encoded>
    </item>
    <item>
      <title>The button was greyed out. That proved nothing.</title>
      <link>https://spoolbound.com/devlog/greyed-out-proved-nothing</link>
      <guid isPermaLink="true">https://spoolbound.com/devlog/greyed-out-proved-nothing</guid>
      <pubDate>Thu, 27 Aug 2026 09:00:00 GMT</pubDate>
      <description>A test asserted that our delete-account button was disabled, and it passed while the code behind the button would happily delete an account on an empty box. Here is the shape of that mistake, and the nine mutations we used to find the rest of them.</description>
      <content:encoded><![CDATA[<p>Last time I wrote about checks that return an honest green about something adjacent to the thing you cared about. This one is the same family, but it is the most expensive member of it, because the check in question guards an irreversible action.</p>
<p>Spoolbound can create an account, so it has to be able to delete one from inside the app. That is a store requirement, and our privacy policy promises it in plain words: you confirm in writing, you do not email anyone, there is no waiting period. Both halves of that sentence are claims about a binary, made in a document anyone can read.</p>
<p>So the question is not "did we build a delete button". It is <strong>how do you know the guard on it is a guard</strong>.</p>
<h2>The mistake, in one line</h2>
<p>The destructive button is disabled until you type the confirmation word. The obvious test asserts exactly that: build the panel, look at the button, assert it is disabled.</p>
<p>That test passes on a panel that will delete your account the moment anything presses the button.</p>
<p><code>disabled</code> suppresses the press signal. So asserting it tells you about the <strong>styling</strong> of the gate, not the gate. Any edit that rebuilds the screen without re-syncing the button, or wires a second control to the same handler, or flips the flag for a "busy" state, removes the protection without touching a line that looks like protection.</p>
<p>The fix is to assert both halves: that the button refuses, <strong>and</strong> that the function behind the button refuses when it is called directly with the field empty.</p>
<h2>Nine mutations</h2>
<p>We do not trust a gate here until we have watched it go red on a real violation. So every guard in that file was broken deliberately, one at a time, and the failures recorded.</p>
<p><strong>M2 is the one the whole file exists for.</strong> Delete the confirmation check from inside the handler, but leave the greying-out alone. Six assertions failed — and <em>"the destructive button is present and disabled" stayed green</em>. That is the entire lesson in one line: the old test would have signed off a panel whose handler deletes an account on an empty box.</p>
<p><strong>M4 is the one that earned the mutation pass its keep.</strong> We broke the code that decides whether the server call succeeded, so a failed deletion would be reported to the player as a success. First run: <strong>everything passed.</strong> Not because the panel was fine, but because a <em>second</em>, unrelated guard further down caught the contradiction and returned before any damage. The file was proving the belt and saying nothing whatsoever about the braces. We closed it by asserting what the mutated panel actually <em>says</em> to the player, which is the thing a person would have noticed and the test had not been looking at.</p>
<p><strong>M7 passed, and that was correct.</strong> We reordered the tabs in the settings panel so Account was second, expecting the "deletion is two taps" assertion to fail. It did not, because the panel explicitly opens on the Account tab regardless of the order the strip was built in. Nothing should have gone red, and nothing did.</p>
<p>That one is worth dwelling on, because the mutation log had previously recorded this line as producing two failures — a description of an edit nobody actually applied. <strong>A mutation log that mis-describes its own edit is the same bug as the test it is supposed to validate</strong>: a record that looks like evidence and is not. The real "deletion becomes three taps" mutation is a different one line, and it fails properly.</p>
<p><strong>M9 was the subtle one.</strong> The test reads the confirmation word from the same constant the panel does, rather than hard-coding it — deliberately, so re-wording the prompt does not break the test. Which raises the obvious worry: empty that constant, and do you disarm the gate and the test at the same time? You do not, and the reason is that one scenario checks the button's state <em>before typing anything at all</em>. The check that saves you is the one that makes no assumptions about the input.</p>
<h2>Two rules that came out of it</h2>
<p><strong>Drive the real thing, not a convenient copy.</strong> The test opens the actual app shell and taps through it, rather than constructing the panel directly. A panel that works when a test builds it and cannot be reached from the running app is not a feature; it is a screenshot. That distinction is not academic — it is a state this project has been in.</p>
<p><strong>Check refusals more than one way.</strong> Every "it did not delete" assertion checks three independent things: no request was sent, the account still exists, and the session is still signed in. Those can only disagree if something is lying, and the point of three is to notice when they do.</p>
<p>And one small courtesy, because a test that costs you something is a test people switch off: the whole thing runs against a fake server with no network in it, at an address under a domain guaranteed never to resolve, and it snapshots the developer's own session and puts it back before reporting. Signing yourself out of your own game is a rude way to discover a test touched something real.</p>
<h2>The thing under all of it</h2>
<p>A promise in a published document is only worth what the binary does. The gap between them is not visible from either side on its own — you can read the policy all day and learn nothing about the code, and read the code all day and never notice it contradicts a sentence somebody published.</p>
<p>A test that presses the real button is the only thing that holds those two together. Which means when the interface changes, as it just did, the <em>test</em> breaks — loudly, in a place someone is looking — instead of the promise quietly becoming false.</p>]]></content:encoded>
    </item>
    <item>
      <title>Three green lights that were not wired to anything</title>
      <link>https://spoolbound.com/devlog/green-lights</link>
      <guid isPermaLink="true">https://spoolbound.com/devlog/green-lights</guid>
      <pubDate>Thu, 27 Aug 2026 09:00:00 GMT</pubDate>
      <description>A check that cannot come back negative is decorative. We have a rule about that, we enforce it on the game, and this week it caught me three times in four days on the website instead.</description>
      <content:encoded><![CDATA[<p>There is a rule on this project that gets quoted more than any other: <strong>a check that cannot come back negative is decorative.</strong> Before you trust a gate, break the thing it protects and confirm it screams.</p>
<p>It is a good rule. It has caught a test suite that ran zero tests and exited zero, a quality gate that passed a deliberately blurred image, and a colour check that sorted its input before testing whether it ascended, so it could never fail.</p>
<p>This week it caught me three times, on the website rather than the game, and all three were the same bug wearing different clothes. Writing them down because the shape is the useful part — I did not recognise the second one as the same mistake as the first, and I had made the first one four days earlier.</p>
<h2>One: the exit code that belonged to the wrong process</h2>
<p>I was mutation-testing a gate. Break the input, confirm it fails, restore. The command was roughly this:</p>
<pre><code>node scripts/check-thing.mjs | tail -3
echo "exit=$?"</code></pre>
<p>It printed <code>exit=0</code> every time. Every mutation reported success. I nearly concluded the gate was broken.</p>
<p><code>$?</code> in a pipeline is the exit status of the <strong>last</strong> command in it. I was reading whether <code>tail</code> had succeeded, and <code>tail</code> always succeeds. The gate underneath was failing correctly and shouting about it; I had built a machine to discard the answer and report on the messenger instead.</p>
<p>The fix is to stop piping while measuring:</p>
<pre><code>node scripts/check-thing.mjs &gt; /dev/null 2&gt;&amp;1
echo "exit=$?"</code></pre>
<p>I hit the same thing later with a <code>grep</code> in the middle of an <code>&amp;&amp;</code> chain — the grep <em>found</em> the failure, so it exited zero, so the <code>&amp;&amp;</code> fired and the next command ran as though everything had passed. A search that succeeds at finding bad news is not the same as bad news being absent, and shell status codes do not distinguish those.</p>
<h2>Two: the flag that only did half its job</h2>
<p>Pages that should not be indexed carry a <code>noindex</code> flag in our route table. I added a page, set the flag, and moved on.</p>
<p>The flag removed the page from the sitemap. That is all it did. It emitted no <code>&lt;meta name="robots"&gt;</code> at all, so a crawler that found the URL any other way — a link, a redirect, someone pasting it — would have indexed it happily.</p>
<p>Worse, the sitemap and the page were now telling a crawler two different things, which is its own class of warning. The flag was not merely incomplete; it was decorative in the precise sense of the rule. It looked like a control. Nothing tested that it controlled anything, because the only observable effect was in a file nobody diffs.</p>
<p>It now emits the meta <em>and</em> filters the sitemap, and the build prints which routes it omitted, so a route silently dropping out is visible instead of mysterious.</p>
<h2>Three: the one that actually mattered</h2>
<p>We publish a document that the app and the site both have to agree with. It has a marker partway down: everything below it is internal working notes and only the part above it is meant to be public.</p>
<p>Our generator did not know about the marker. It published the whole file.</p>
<p>That is a plain bug and it would be a short entry, except for what I did next. I checked. I searched the live page for a string I knew was in the private half, got no match, and concluded the marker was being honoured.</p>
<p>The string I searched for had been added to the document that same day. It had never been in the version the live site was built from. <strong>My probe could not have found anything, and I read that as proof of absence.</strong> A single-string check, treated as evidence of a general property, on the one occasion the general property was false.</p>
<p>It had been wrong for days and my check had been incapable of telling me.</p>
<h2>What the three have in common</h2>
<p>None of them was a wrong answer. All three were <strong>a correct answer to a question I had not asked</strong>.</p>
<ul>
<li>Did <code>tail</code> run? Yes. That was never in doubt.</li>
<li>Is the page in the sitemap? No. Also not the question.</li>
<li>Is this one specific string on the page? No. Still not the question.</li>
</ul>
<p>Every one of those returned an honest green, and every green was about something adjacent to the thing I cared about. The failure mode is not a broken check — a broken check gets noticed. It is a <strong>working check pointed slightly to the left of the problem</strong>, which is invisible precisely because it works.</p>
<p>The corollary we already had was: mutation-test the gate. The one I would add is narrower and I would have found all three faster with it:</p>
<p>&gt; <strong>Before trusting a green, write down the exact question you think it answered — then read the code and check that is the question it asked.</strong></p>
<p>The <code>tail</code> did not answer "did the gate fail". The sitemap did not answer "will this be indexed". The grep did not answer "is the private half leaking".</p>
<p>Three greens. None of them wired to the thing I was watching.</p>]]></content:encoded>
    </item>
    <item>
      <title>Which way is front?</title>
      <link>https://spoolbound.com/devlog/which-way-is-front</link>
      <guid isPermaLink="true">https://spoolbound.com/devlog/which-way-is-front</guid>
      <pubDate>Mon, 24 Aug 2026 09:00:00 GMT</pubDate>
      <description>Every printer needs a clear tile in front of it. Until this week "front" was the same fixed direction for every machine in the shop, which made clearance a tax you paid rather than a space you could plan.</description>
      <content:encoded><![CDATA[<p>The request was one sentence: make the printers behave like physical objects.</p>
<p>That sounds like an art note. It turned out to be a simulation change, a save-format decision, and a long afternoon spent on where the light falls.</p>
<h2>The rule that was already there</h2>
<p>Clearance shipped yesterday. Every printer needs one empty tile in front of it, for the same reason a real machine does — you have to open it and reach into it, and a machine you have shoved against three neighbours is a machine you cannot work on.</p>
<p>Good rule. One problem: <strong>"front" was a fixed convention.</strong> It was always the row nearest the camera, for every machine in the game, because nothing in the model had ever carried a direction. Not the printers, not the benches. Not even your own character.</p>
<p>So every machine faced the same way, and their clear tiles could never be shared. Two printers side by side each carved out a private empty strip in front of themselves. You could not point them at a common aisle, because there was no pointing.</p>
<p>That makes clearance a tax: you pay a tile per machine and get nothing back for arranging them well. Rotation is the difference between that and a layout puzzle — turn a row of machines to face one corridor and the empty space starts doing two jobs at once.</p>
<h2>The save version we deliberately did not bump</h2>
<p>The specification for this change said to bump the save format version and add a documented no-op rung to the migration ladder. We didn't, and the reasoning is worth writing down.</p>
<p><code>facing</code> is an <strong>optional</strong> field, and its absence already means the one direction every machine used to have. An old save is therefore already a valid new save — nothing about its shape changed, so there is nothing for the ladder to carry. A version rung that exists only to announce that nothing happened is noise, and the next person to read that ladder has to spend attention ruling it out.</p>
<p>The evidence that the default really is the old behaviour, rather than something we asserted: <strong>all 252 existing tests passed unmodified.</strong> Not adjusted, not re-recorded. If a later change ever makes <code>facing</code> required, <em>that</em> is the change that earns the bump — and it is written down where the decision lives, so nobody has to rediscover the reasoning.</p>
<h2>The light was the hard part</h2>
<p>The art has a rule we call the value-band law: the key light is fixed to the top-left <strong>of the screen</strong>, and floors, outlines, machine bodies and accents each occupy a locked brightness range. Machines read light-against-dark by construction rather than by somebody's judgement on the day.</p>
<p>Here is what that collides with. When you draw a box in this projection, only two of its side faces are ever visible: the two that meet the corner nearest the camera. The old code found them with a hardcoded corner index, which was correct for a slightly embarrassing reason — every call site happened to build its rectangle the same way round, so the nearest corner was always the same index.</p>
<p>Turn the box and that stops being true. The four corners move, but their <strong>cyclic</strong> order survives: it is still a rectangle walked round in sequence. So the visible pair is still "whichever two meet the nearest corner" — that corner is just no longer the index it used to be.</p>
<p>The four cases got worked through by hand before any code was written. Then, deliberately, they were <strong>not</strong> implemented as the formula that derivation produced. The code instead takes an argmax over the points it has already projected to the screen.</p>
<p>Those two are the same arithmetic when they are right, and they fail completely differently. A sign error in the formula draws a face on the object's far side — where nobody can see it, because it is facing away. The same sign error in the argmax version gives you a <strong>visibly wrongly-lit box</strong>, immediately, on the first frame.</p>
<p>That is this project's house rule pointed at geometry. When two correct-looking implementations are available, take the one whose failure mode is loud.</p>
<p>The same relabelling feeds the top face's shading and the corners the box hands back, so a decoration authored against the front — the enclosed machine's door, the insulation bands up the side of the tall one — follows the machine round automatically instead of staying on whichever side was originally west.</p>
<h2>What is not done</h2>
<p>Nothing in the game lets you turn anything yet.</p>
<p>The model carries facing, the renderer honours it, the depth sorting and the little ready badges above each machine were all corrected to use the turned footprint rather than the original one. All of that last part is a <strong>no-op today</strong> — it changes nothing until a machine is actually turned — and it was done now so that the day the control lands, it is already right, instead of a second pass rediscovering the same bug.</p>
<p>So the feature is currently real, correct, and invisible. That is a strange thing to ship and a normal thing to build.</p>
<p>The tile in front of a printer was always going to cost you something. What changes is whether it is only a cost.</p>]]></content:encoded>
    </item>
    <item>
      <title>How do you test your tests?</title>
      <link>https://spoolbound.com/devlog/how-do-you-test-your-tests</link>
      <guid isPermaLink="true">https://spoolbound.com/devlog/how-do-you-test-your-tests</guid>
      <pubDate>Mon, 24 Aug 2026 09:00:00 GMT</pubDate>
      <description>Every safety check in this project has been watched failing on a real problem before its passing result was believed, because a check that cannot come back negative is indistinguishable from one that works.</description>
      <content:encoded><![CDATA[<p>The house rule here is one sentence:</p>
<p>&gt; <strong>A check that cannot come back negative is decorative.</strong></p>
<p>Before we trust any gate, we break the thing it protects and confirm it screams. Then we put the thing back. It costs a couple of minutes and it has never once been wasted.</p>
<p>That sounds like paranoia until you see the list.</p>
<h2>What the rule has caught, in this project alone</h2>
<p><strong>A test suite that ran zero tests and reported success.</strong> Exit code 0, green output, nothing executed. It had been "passing" for a while.</p>
<p><strong>A quality gate that passed a deliberately blurred image.</strong> The gate existed specifically to catch blurry output. We fed it something obviously wrong to watch it complain, and it congratulated us.</p>
<p><strong>A check for a rising difficulty curve that sorted its input before testing whether it rose.</strong> Read that twice — it sorted the numbers ascending, then verified they ascended. It could not fail. It would have gone on not failing forever, through any amount of broken tuning.</p>
<p><strong>A config comment claiming an audio problem was solved</strong> when nothing had been done about it at all. Not a check, but the same failure: a confident green signal with nothing behind it.</p>
<p>None of those were visible by reading the code. Every one looked correct. They were found by making something run and watching what it actually did.</p>
<h2>The corollary, which is sharper</h2>
<p>We added a second rule after a near-miss:</p>
<p>&gt; <strong>A mutation test that does not confirm the mutation landed is itself decorative.</strong></p>
<p>Here is how that one nearly got us. To check a gate, you deliberately break the thing it guards — change a value, run the gate, watch it go red. Standard.</p>
<p>We tried to break a price on purpose. The edit silently matched nothing: the string we were searching for had moved. So the "mutated" run was really an unmutated run, the gate reported green, and green was the correct answer to the question we accidentally asked.</p>
<p><strong>A failed experiment and a well-protected system produce identical output.</strong> The only thing separating them is whether you checked that your sabotage actually took. Now every probe asserts the mutation landed before it believes the result.</p>
<h2>Why this is worth the two minutes</h2>
<p>The value of a check is not that it passes. It is that it <em>could</em> fail and didn't. A check that cannot fail carries no information at all — but it looks exactly like one that carries a lot, and it sits in your build going green while the thing it was written to protect quietly rots.</p>
<p>The failure mode is not "the check is wrong". It is "the check is absent, and you believe it is present." Those are very different situations to be in, and from the outside they are the same colour.</p>
<p>So: break it, watch it scream, put it back. Every time. If it does not scream, you have just learned something considerably more important than whether the code works.</p>]]></content:encoded>
    </item>
    <item>
      <title>We tried to fix a boring stretch with more content, four times. The data said stop.</title>
      <link>https://spoolbound.com/devlog/four-fillers</link>
      <guid isPermaLink="true">https://spoolbound.com/devlog/four-fillers</guid>
      <pubDate>Mon, 24 Aug 2026 09:00:00 GMT</pubDate>
      <description>There was a stretch of the game with nothing meaningful to buy. Four different fillers were built and measured, and all four failed in exactly the same way — by making the run longer instead of fuller.</description>
      <content:encoded><![CDATA[<p>The obvious fix for a boring stretch is more content. Put something in the gap. We believed that for long enough to build four different things and measure every one of them.</p>
<p>All four failed. Not narrowly, and not for four different reasons — they failed <em>identically</em>, which is what eventually made the problem legible.</p>
<h2>The gap</h2>
<p>Somewhere in the middle of a run there was a stretch with nothing meaningful to buy. Not idle time — you were still working, still collecting, still keeping machines fed. But every decision had already been made. You were waiting on a number to get big enough, and nothing you did changed which number or how big.</p>
<p>That is the worst combination available: work without decisions.</p>
<h2>The four attempts</h2>
<p>Each one was a real thing, built and measured against the same probe, from the same cold start:</p>
<p>1. <strong>A cheap mid-tier machine</strong> to sit between the ones that bracket the gap. 2. <strong>A consumable upgrade</strong> — something to spend on repeatedly rather than once. 3. <strong>A cosmetic tier</strong>, on the theory that spending is spending. 4. <strong>A cheaper version of the milestone itself</strong>, to bring the far wall nearer.</p>
<p>Four different shapes, four different price points. The runs came back looking the same as each other and worse than the baseline.</p>
<h2>Why they all failed the same way</h2>
<p>The arithmetic is embarrassing once you see it.</p>
<p><strong>Every dollar spent on something new pushes the next real milestone further away by exactly the time it takes to earn that dollar back.</strong></p>
<p>That is not a tuning problem. It is a conservation law. Money in the game is time, so anything you add to the middle of the run is paid for out of the end of it. The gap does not get filled — it gets <em>moved</em>, and the run gets longer while it happens.</p>
<p>Measured against the baseline, the fillers extended the run by <strong>up to 15%</strong>. We had built four things, each of which made the game take longer to reach the same place, and each of which felt like progress while we were building it.</p>
<p>The cosmetic tier is the interesting one, because it should have escaped the trap — cosmetics do not compete with progression. It failed anyway. If a player spends on it, the money is still gone. If they don't, it isn't content, it's a menu.</p>
<h2>What actually worked</h2>
<p>The wall could not be removed. So we stopped trying to remove it and made it <strong>legible</strong>.</p>
<p>Every unaffordable thing now states its real requirement, in plain language, on the thing itself:</p>
<pre><code>Computer L3
All 2 installed
Research: Print-in-Place</code></pre>
<p>— along with an estimate of when you will be able to afford it at your current rate.</p>
<p>Nothing was added. Nothing got cheaper. The gap is exactly as long as it was. But the stretch stopped being a wait and became a <em>plan</em>, because you can see what you are waiting for and roughly how long it will take. A player who knows they are eleven minutes from the thing they want is not bored in the same way as a player who knows only that they are not there yet.</p>
<h2>The part worth stealing</h2>
<p>We spent four builds on the first idea we had, because it was obvious and because each attempt felt like work. What broke the loop was not a better idea. It was measuring the fourth failure and noticing it had the same shape as the first three.</p>
<p><strong>Four failures that look identical are not four failures. They are one result.</strong> If we had lined them up sooner we would have found the conservation law sooner, and built one thing instead of five.</p>
<p>The instinct to fix a content problem with content is strong enough that it survived three refutations. It is worth asking, early, whether the thing you keep failing to fix is a thing that can be fixed at all — or whether the honest move is to stop hiding it and start explaining it.</p>]]></content:encoded>
    </item>
    <item>
      <title>Three layers of save protection, defeated by one missing try/catch</title>
      <link>https://spoolbound.com/devlog/three-layers-one-missing-catch</link>
      <guid isPermaLink="true">https://spoolbound.com/devlog/three-layers-one-missing-catch</guid>
      <pubDate>Mon, 24 Aug 2026 09:00:00 GMT</pubDate>
      <description>A single corrupt value could destroy a save, and every safety net we had built failed at once — silently, and in a way that was invisible by reading the code.</description>
      <content:encoded><![CDATA[<p>This one is unflattering, so here it is in full.</p>
<p>A save could be destroyed by a single bad value. Not corrupted — <em>destroyed</em>, overwritten by a brand new empty shop. And the three independent safety systems built specifically to stop that all failed at the same moment, without printing a single warning.</p>
<h2>What was supposed to happen</h2>
<p>The save system has three layers, and on paper they cover each other.</p>
<p>1. <strong>The current save</strong>, written as you play. 2. <strong>A backup</strong>, kept one step behind, so a bad write has something to fall back to. 3. <strong>A quarantine</strong>, which takes anything unreadable and stores it verbatim instead of deleting it — on the theory that a save we cannot parse today might be recoverable tomorrow, and is certainly worth more than nothing.</p>
<p>On top of that sits a <strong>migration ladder</strong>: a chain of small upgrades that walks an old save forward, version by version, until it matches what the current build expects.</p>
<p>Before the ladder runs, a sanity check looks at the data. That check deliberately only validates fields that have existed in <strong>every</strong> version of the format. It has to. A check that demanded today's fields would reject the very old saves the ladder exists to carry forward — it would be a migration system that refuses to migrate anything.</p>
<h2>What actually happened</h2>
<p>That deliberate narrowness is the hole.</p>
<p>A field introduced in a <em>later</em> version could hold a corrupt value. The sanity check does not look at it, correctly, because it cannot. So the save passes the gate and enters the ladder. Then a migration step touches that field and throws.</p>
<p>And the throw happened <strong>inside</strong> the migration, which is to say inside the loader, which is to say it escaped the loader entirely. From there, everything downstream did the wrong thing for a defensible reason:</p>
<ul>
<li>The <strong>quarantine</strong> never ran, because quarantine is for data that fails to parse. This data parsed fine. It failed later, somewhere the quarantine was never wired to hear about.</li>
<li>The <strong>backup</strong> was never tried, because nothing in the crash path knew a load had been attempted and failed. The exception went straight past it.</li>
<li>The <strong>boot code</strong> caught the error at the top and found itself holding no save. It could not distinguish <em>"there is a damaged save here"</em> from <em>"this is a new player"</em>, and those two situations have opposite correct responses. It picked the wrong one: it started a fresh shop — and wrote it straight over the top of the save it had just failed to read.</li>
</ul>
<p>Three layers, one uncaught exception, and the failure mode is the exact thing all three exist to prevent.</p>
<h2>The comment that was wrong</h2>
<p>The module's own documentation stated plainly that a save is never destroyed by this code.</p>
<p>That was not a lie anybody told. It was true for every failure the author had thought of, and false for the one they had not — which is the ordinary way a comment goes wrong. Documentation describes the cases you imagined. It cannot describe the case you missed, and it will keep confidently asserting the general claim long after the general claim stops holding.</p>
<p>We now treat a comment that promises a guarantee as a <strong>claim requiring a test</strong>, not as documentation.</p>
<h2>The fix, and proving it</h2>
<p>The fix is unglamorous: the migration ladder runs inside a boundary that treats <em>any</em> failure — not just a parse failure — as "this save is damaged", which routes it to the quarantine and then to the backup, and the boot path can now tell a damaged save from an absent one.</p>
<p>The part that matters is the order it was done in. <strong>The original destruction was reproduced first.</strong> A test was written that fed the loader the exact shape of bad data, and it was watched destroying the save — confirming the failure was real, that the test could see it, and that the test would have caught it. Only then was the fix applied, and the same test watched going green.</p>
<p>That order is not a formality. A test written after a fix, against code that already works, proves only that the code does what it currently does. It never demonstrates that it can detect the bug, and a test that cannot fail on the thing it is named after is decorative.</p>
<p>This was caught in development, before anybody outside the project had a save to lose. That is luck as much as process — but the process is what turned it from a mystery into a reproducible case in an afternoon.</p>]]></content:encoded>
    </item>
    <item>
      <title>Your pixel-art game cannot have smooth zoom. Here is the compromise we shipped.</title>
      <link>https://spoolbound.com/devlog/smooth-zoom-sharp-pixels</link>
      <guid isPermaLink="true">https://spoolbound.com/devlog/smooth-zoom-sharp-pixels</guid>
      <pubDate>Sun, 23 Aug 2026 09:00:00 GMT</pubDate>
      <description>Pinch-to-zoom ran straight into a physical constraint rather than a taste question — crisp pixels and fluid scaling are, at the naive level, mutually exclusive. The fix was to stop treating it as one decision.</description>
      <content:encoded><![CDATA[<p>The bug report was two sentences. The screen went black while pinch-zooming on a real phone, and then, once that was fixed: could the zoom feel fluid instead of locked into some positions?</p>
<p>The first half was a bug. The second half was not, and it took a while to work out why.</p>
<h2>Why sharpness and smoothness fight</h2>
<p>The game draws at a <strong>low internal resolution</strong> and lets the device scale that canvas up to its real pixel count. That is the standard way to make pixel art look like pixel art on a screen with four times too many pixels, and it has one hard condition attached: it is only crisp at <strong>whole-number scale factors</strong>.</p>
<p>At 2×, every art pixel becomes exactly four screen pixels. Clean edges, no ambiguity.</p>
<p>At 1.5×, an art pixel has to cover one and a half screen pixels, and there is no such thing as half a pixel. The renderer has to share — some art pixels get two screen pixels, some get one — and the result is soft, uneven, and it <em>shimmers</em> as you move, because which pixels get doubled changes from frame to frame.</p>
<p>So the two things being asked for genuinely conflict. Continuous zoom means arbitrary scale factors. Arbitrary scale factors mean fractional ones. Fractional ones mean soft, crawling art.</p>
<h2>Not assuming which way to break the tie</h2>
<p>There were two honest options and no correct one, so it went back as a question rather than a decision:</p>
<ul>
<li><strong>Continuous, always.</strong> The zoom feels perfect and the art is very slightly soft at most zoom levels, all the time.</li>
<li><strong>Smooth while your fingers are moving, then snap crisp when you let go.</strong> Fluid during the gesture; resolves to a whole-number step the instant you lift.</li>
</ul>
<p>The second one was chosen, and that is what shipped. While the gesture is live, the view scales fluidly in real time. The moment you lift your fingers, it settles onto the nearest crisp step.</p>
<p>It is worth being clear that this is a compromise and not a trick. There is a fraction of a second where the art is soft. It happens to be the fraction of a second where your fingers are covering the screen and everything is moving, which is exactly when you are least able to notice.</p>
<h2>Two details worth the extra effort</h2>
<p><strong>The snap picks the nearest step by ratio, not by difference.</strong> Zoom is multiplicative — going from 1× to 2× is the same <em>kind</em> of change as going from 2× to 4×. So the fair midpoint between 1× and 2× is not 1.5×, it is their geometric mean, about <strong>1.41×</strong>. Snapping on plain arithmetic distance makes the step feel biased toward the lower end, and it feels wrong in a way that is very hard to name if you do not know what you are looking for.</p>
<p><strong>Pinching past the end of the range rubber-bands.</strong> Rather than hitting a dead wall at maximum zoom, the view keeps responding with increasing resistance and springs back when you let go, the way a scroll view does at the top of a list. A hard stop reads as a bug. Resistance reads as a boundary.</p>
<h2>What the whole thing is actually about</h2>
<p>The instinct when someone says "this should feel smoother" is to make it smoother. Here that would have meant quietly accepting permanently soft art in a game whose entire visual identity is crisp pixels, to fix a complaint about a gesture that lasts under a second.</p>
<p>The constraint was real and could not be removed. What could change was <strong>when the player pays for it</strong> — and moving that cost into the half-second where nobody is looking turned out to be worth more than solving the problem would have been.</p>]]></content:encoded>
    </item>
    <item>
      <title>We wrote a robot to play our game perfectly, and it found a 55-minute hole</title>
      <link>https://spoolbound.com/devlog/we-measured-our-own-game</link>
      <guid isPermaLink="true">https://spoolbound.com/devlog/we-measured-our-own-game</guid>
      <pubDate>Sat, 22 Aug 2026 09:00:00 GMT</pubDate>
      <description>A probe that plays the real simulation as a flawless player turned out to be the most useful thing in the repository, because it measures the one thing a developer cannot feel.</description>
      <content:encoded><![CDATA[<p>You cannot playtest your own game for pacing. You know where everything is, you know what to buy next, and you have played the first ten minutes four hundred times. Every instinct you have about whether the middle drags is contaminated.</p>
<p>So we stopped guessing and wrote a probe.</p>
<h2>What it does</h2>
<p>It is not a bot in the game. It imports the actual simulation — the same module the app runs, with no rendering, no input and no clock — and plays it as a perfect player. It never misses a collection, never hesitates over a purchase, never wanders off. It walks the whole run from the first dollar to the warehouse and prints what happened.</p>
<p>That last property is the important one. <strong>A perfect player is a floor.</strong> Every gap the probe reports is the shortest that gap can possibly be, because a real person is slower and never faster. If the robot is bored, everybody is bored.</p>
<h2>What it found</h2>
<p>The first run was not close.</p>
<pre><code> 4.0m  printer #2     7.3m  printer #3    10.7m  printer #4
15.0m  printer #5    20.3m  printer #6    75.7m  WAREHOUSE
LONGEST DEAD STRETCH: 55.3 min</code></pre>
<p>Six things to buy across seventy-six minutes, and then a <strong>fifty-five minute stretch with nothing to decide at all</strong>. Not idle time — the tray cap meant the player still had to tap roughly three hundred and forty times inside that stretch simply to stop the machines stalling. Work without decisions, which is the worst combination available.</p>
<p>The opening, meanwhile, was fine. A purchase every three to five minutes, gaps widening gently, a start with no money that teaches the whole loop in about thirty seconds without a tutorial. Nobody had touched that curve and nobody should.</p>
<p>So the problem was not tuning. Cutting the price of the warehouse would have shortened the hole without putting anything in it. <strong>It was a content gap wearing a pacing gap's clothes</strong>, and the two have completely different fixes.</p>
<h2>What we did</h2>
<p>Built the thing that belongs in the middle: a research spine and a set of workstations, so the stretch between the sixth printer and the warehouse has decisions in it rather than distance.</p>
<p>Then re-ran the same probe, because a fix nobody measured is a hope.</p>
<pre><code>lane      warehouse   longest gap   purchases
chamber     100.0m        10.3m         31
freeAir      91.4m        10.8m         31</code></pre>
<p>Six purchases became thirty-one. The longest stretch with nothing to buy went from fifty-five minutes to about ten. The run got <em>longer</em> and stopped being empty, which is the direction you want — a game is not improved by being shorter, it is improved by having something in it.</p>
<h2>The part worth stealing</h2>
<p>The probe is an <strong>instrument</strong>, not a test. It has no assertions of its own: it reports numbers, and a separate gate decides what they mean. That separation is the whole design. A pacing check that goes red on a threshold somebody invented is a threshold somebody invented, enforced forever — so the threshold lives outside the instrument, written down where it can be argued with.</p>
<p>The other half is a rule we keep relearning: <strong>a check that cannot come back negative is decorative.</strong> Before trusting any gate here we break the thing it protects and confirm it screams. That habit has caught a test suite that ran zero tests and exited zero, a quality gate that passed a deliberately blurred image, and a ramp check that sorted its input before testing whether it ascended — so it could never fail.</p>
<p>None of those were visible by reading the code. All of them were found by making something run and watching what it did.</p>]]></content:encoded>
    </item>
  </channel>
</rss>
