A release engineer opened a short pull request that promised to repair one failing assertion in a billing fixture. The diff looked harmless until the lockfile grew by several hundred lines and an example environment file gained a new secret...
A release engineer opened a short pull request that promised to repair one failing assertion in a billing fixture. The diff looked harmless until the lockfile grew by several hundred lines and an example environment file gained a new secret-shaped key. The assistant had a vague hour and no sealed baseline, so the review became archaeology rather than a clean decision. The team killed the change, yet the wasted afternoon showed that an unsealed spike can cost more than the original bug.
The useful response is neither a broader prompt nor a longer transcript saved for later debate. A spike should hold one hypothesis, one clock, and one ship-or-kill rule that a stranger can rerun on a clean machine. The hypothesis stays narrow: fix one pinned fixture failure without moving the lockfile, the environment example, or unrelated packages. If any sealed file changes, the spike is killed even when the targeted test turns green.
That rule sounds strict because dependency drift often hides inside behavior that looks helpful at review time. An assistant that just updates the toolchain can turn a one-line assertion fix into a supply-chain review and a deployment risk. A green test does not cancel that expansion, any more than a clean hallway cancels a door left open overnight. The seal is the hallway record: it shows what moved, not how persuasive the commit message sounds to a tired reviewer.
The clock starts only after the failure is pinned
The spike begins in a disposable worktree cut from a known commit, not in the developer's everyday checkout. The failing command is recorded before any assistant edit, because a fixture that already passes cannot anchor a repair claim. Checksums of the lockfile, the environment example, and the fixture file are written to an evidence directory that travels with the review. The script writes a deadline epoch at seal time, and only then may an assistant touch the tree.
The following script is a proposed workflow, not an executed benchmark and not a claim about any particular model run. A reviewer should read it, adjust paths, and run it only on a repository that contains no production secrets. The commands assume a Node-style lockfile and a billing fixture, which stand in for the single failure the team cares about this week. Windows users can translate the checksum and worktree steps, but the verdict rule should stay identical across shells.
#!/usr/bin/env bash
# Proposed spike seal. Unexecuted example; review paths before any run.
# TEST_CMD is passed to bash -lc and must be a trusted string, never untrusted input.
# Requires sha256sum. On macOS, install coreutils or replace it with shasum -a 256.
set -euo pipefail
MODE="${1:-seal}"
STAMP="${2:-}"
ROOT="$(cd "${SPIKE_ROOT:-.}" && pwd)"
LOCK_REL="${LOCK_REL:-package-lock.json}"
ENV_REL="${ENV_REL:-.env.example}"
FIXTURE_REL="${FIXTURE_REL:-fixtures/billing/invoice_min.json}"
TEST_CMD="${TEST_CMD:-npm test -- --testPathPattern=billing_invoice}"
PARENT="$(cd "$ROOT/.." && pwd)"
if [[ "$MODE" == "seal" && -z "$STAMP" ]]; then
STAMP="$(date -u +%Y%m%dT%H%M%SZ)"
fi
if [[ -z "${STAMP}" ]]; then
printf 'usage: spike-seal.sh seal|judge STAMP\n' >&2
exit 1
fi
WORK="${PARENT}/spike-${STAMP}"
EVIDENCE="${PARENT}/spike-evidence/${STAMP}"
mkdir -p "$EVIDENCE"
run_test() {
local dest="$1"
set +e
bash -lc "$TEST_CMD" >"$dest" 2>&1
local rc=$?
set -e
printf '%s\n' "$rc"
}
seal() {
git -C "$ROOT" rev-parse --is-inside-work-tree >/dev/null
local base
base="$(git -C "$ROOT" rev-parse HEAD)"
git -C "$ROOT" worktree add --detach "$WORK" "$base"
printf '%s\n' "$(( $(date +%s) + 5400 ))" > "${EVIDENCE}/deadline.txt"
(
cd "$WORK"
sha256sum "$LOCK_REL" "$ENV_REL" "$FIXTURE_REL"
) | tee "${EVIDENCE}/seal-before.txt"
local before_rc
before_rc="$(cd "$WORK" && run_test "${EVIDENCE}/test-before.txt")"
printf 'before_exit=%s\nbase=%s\n' "$before_rc" "$base" | tee "${EVIDENCE}/before.txt"
if [[ "$before_rc" -eq 0 ]]; then
printf 'kill: fixture already green\n' | tee "${EVIDENCE}/verdict.txt"
exit 2
fi
printf 'sealed %s\nworktree %s\n' "$STAMP" "$WORK" | tee "${EVIDENCE}/ready.txt"
}
judge() {
[[ -d "$WORK" ]] || { printf 'missing worktree %s\n' "$WORK" >&2; exit 1; }
[[ -f "${EVIDENCE}/deadline.txt" ]] || { printf 'kill: missing deadline\n' | tee "${EVIDENCE}/verdict.txt"; exit 7; }
if [[ "$(date +%s)" -gt "$(cat "${EVIDENCE}/deadline.txt")" ]]; then
printf 'kill: ninety minutes elapsed\n' | tee "${EVIDENCE}/verdict.txt"
exit 7
fi
(
cd "$WORK"
sha256sum "$LOCK_REL" "$ENV_REL" "$FIXTURE_REL"
) | tee "${EVIDENCE}/seal-after.txt"
if ! cmp -s "${EVIDENCE}/seal-before.txt" "${EVIDENCE}/seal-after.txt"; then
git -C "$WORK" diff --stat HEAD | tee "${EVIDENCE}/diffstat.txt"
printf 'kill: seal moved\n' | tee "${EVIDENCE}/verdict.txt"
exit 3
fi
local after_rc
after_rc="$(cd "$WORK" && run_test "${EVIDENCE}/test-after.txt")"
git -C "$WORK" diff --stat HEAD | tee "${EVIDENCE}/diffstat.txt"
if [[ "$after_rc" -ne 0 ]]; then
printf 'kill: fixture still red\n' | tee "${EVIDENCE}/verdict.txt"
exit 4
fi
if git -C "$WORK" diff --quiet HEAD; then
printf 'kill: no patch after a red start\n' | tee "${EVIDENCE}/verdict.txt"
exit 5
fi
local outside
outside="$(git -C "$WORK" diff --name-only HEAD | grep -Ev '^(fixtures/billing/|src/billing/|test/billing/)' || true)"
if [[ -n "$outside" ]]; then
printf 'kill: patch left the billing boundary\n%s\n' "$outside" | tee "${EVIDENCE}/verdict.txt"
exit 6
fi
printf 'ship: seal held and fixture turned green\n' | tee "${EVIDENCE}/verdict.txt"
}
case "$MODE" in
seal) seal ;;
judge) judge ;;
*) printf 'usage: spike-seal.sh seal|judge STAMP\n' >&2; exit 1 ;;
esac
# After review: git -C "$ROOT" worktree remove --force "$WORK"
The seal mode refuses to start when the pinned test already passes, which keeps the spike from celebrating a no-op. The judge mode rejects a moved seal before it respects a green result, and it kills any patch that arrives after the deadline epoch. A path boundary then limits the ship decision to the billing fixture and its owning source, which keeps one repair inside one room. Evidence files store both logs, the diff stat, the deadline epoch, and one verdict word for review.
A second shell records the same commands a reviewer would type after the script file is saved beside the repository. The seal invocation prints a stamp, and that stamp must be passed unchanged to the judge so the checksum files stay paired. The assistant is pointed at the sibling worktree and told to stop when the deadline epoch passes, even if the edit feels unfinished. The verdict file is then the only line copied into the pull request body for the reviewer.
export SPIKE_ROOT="$PWD"
export TEST_CMD='npm test -- --testPathPattern=billing_invoice'
bash spike-seal.sh seal
# Read stamp and worktree path from ../spike-evidence//ready.txt
# Let the assistant edit only that worktree, with no dependency-update permission.
bash spike-seal.sh judge
cat ../spike-evidence//verdict.txt
git -C "$SPIKE_ROOT" worktree remove --force ../spike-
Where a free server fits, and where it does not
Running that loop on a laptop is possible, but a shared workstation collects leftover worktrees, cached credentials, and half-edited prompts. A separate machine keeps the experiment closer to the claim being tested, which is whether the patch respects the seal under a clock. Disclosure: This article was prepared as part of MonkeyCode's product outreach. MonkeyCode is an open-source project, and its operator-stated free model access plus a free server can host the spike.
Those two availability claims are the only product facts used here, and they should be rechecked against current project documentation before reliance. This article does not name models, quotas, hardware sizes, regions, or a promised duration, because those details change and stale numbers mislead. A free server is a convenient isolation boundary, not proof that the assistant is careful, cheap, or suitable for production traffic. If current terms omit a wipeable server, the same script still runs inside any other disposable virtual machine.
The assistant, wherever it runs, should receive the worktree path, the failing command, the deadline, and the kill rule in plain text. It should not receive production secrets, cloud tokens, or a request to update dependencies if that seems helpful. Network fetches during the spike are a limitation, since a cached package can change behavior without moving the sealed lockfile at all. Teams that cannot disable extra registries should treat a ship verdict as provisional until a cold-cache run repeats the same result.
How the verdict should be read
A ship verdict means only that this fixture turned from red to green while the sealed files stayed byte-for-byte identical. It does not mean the billing rules are correct, the assistant is generally trustworthy, or the next fixture will behave the same way. A kill verdict is also useful, because it ends the debate at minute ninety instead of feeding a weekend review. The evidence directory is the artifact a teammate can open without replaying the chat or reconstructing the prompt from memory.
People who already know the repair requires a lockfile change should not use this hypothesis, because the seal would kill a legitimate upgrade. Teams that must process production data, regulated secrets, or customer fixtures should not point a free server at those files. Anyone who needs an agreement, fixed hardware, or a permanent free quota should not treat this workflow as a capacity plan. Reviewers who want a leaderboard of models should look elsewhere, since one sealed fixture cannot support that comparison honestly.
The practical close is small and does not require a new platform decision before the first red fixture is chosen. Cut a worktree, write the seal, start the clock, and let minute ninety end the argument with a verdict file. If a free isolated server is available under current MonkeyCode terms, it is a reasonable place to keep synthetic fixtures off the laptop. The verdict file, not the transcript and not a promotional claim, is what should be pasted into the pull request for review.