At 9:10 the lab is already warm. A student has a green test line from an agent run and a finger resting on the merge control, the way a press operator rests a hand on the cycle button before the die has been proved. The instructor sets a sc...
At 9:10 the lab is already warm. A student has a green test line from an agent run and a finger resting on the merge control, the way a press operator rests a hand on the cycle button before the die has been proved. The instructor sets a scratched plate on the bench. Nothing that looks like production steel gets cut until that plate says the die is true.
This workshop runs ninety minutes. Every student leaves with a checker that reruns on a clean laptop, against a directory of ordinary files. The class does not grade the agent's prose. It grades whether a small change survived a fixed plate: allowed files only, tests actually executed, a stop reason written down, and no secret-shaped string in the diff.
Free model access makes the first attempt cheap. A cheap attempt is also how an unreviewed patch reaches a shared branch. The plate exists so the cheap attempt still has to pass a gauge.
Disclosure: This article was prepared as part of MonkeyCode's product outreach. MonkeyCode is the open-source project this bench points at, because the operator states that free model access and a free server option are available. The handout does not bake in a token count, a model name, or a rental duration. Those figures move, and a number copied from an old post is how a lab lies to itself.
Ten minutes before the room starts, the instructor opens the current project documentation and writes the live limit on the whiteboard. If the page and a slide disagree, the page wins. Students who cannot see that page do not guess. They skip the hosted run and still complete the local plate.
The plate is a tiny fixture, not a tour of a framework. Students receive two files that already pass, plus one file the agent may touch. The task is dull on purpose. A function clamp_span should reject an inverted range and return the clipped length. A tryout plate that already contains a clever bug teaches the die, not the checker.
def clamp_span(start: int, end: int, limit: int) -> int:
if limit int:
root = Path(run_dir)
report = json.loads((root / "report.json").read_text())
patch = (root / "diff.patch").read_text()
reasons = []
if report.get("tests_exit") != 0:
reasons.append("tests_failed")
if report.get("tests_ran") is not True:
reasons.append("tests_not_executed")
if not report.get("stop_reason"):
reasons.append("missing_stop_reason")
extra = set(report.get("files_touched", [])) - ALLOWED
if extra:
reasons.append("off_plate:" + ",".join(sorted(extra)))
if SECRET.search(patch):
reasons.append("secret_shaped_text")
changed = sum(
1 for line in patch.splitlines()
if line.startswith("+") or line.startswith("-")
)
if changed > MAX_LINES:
reasons.append("diff_over_plate")
verdict = "pass" if not reasons else "fail"
print(json.dumps({
"verdict": verdict,
"reasons": reasons,
"changed_lines": changed,
}))
return 0 if verdict == "pass" else 1
if __name__ == "__main__":
sys.exit(main(sys.argv[1] if len(sys.argv) > 1 else "run"))
A passing report is short. The JSON below is a fixture, not a log captured from a live server. Students replace it with their own export before they trust the exit code.
{
"tests_ran": true,
"tests_exit": 0,
"stop_reason": "tests_green_and_allowlist_held",
"files_touched": ["clamp_span.py"]
}
The matching patch is also a fixture. It is small enough to read aloud. The change is plausible, which is exactly why the plate must not be asked to bless it alone.
--- a/clamp_span.py
+++ b/clamp_span.py
@@ -1,4 +1,4 @@
- return min(end, start + limit) - start
+ return max(0, min(end, start + limit) - start)
The command line is the whole ceremony. Two copies, one process, one exit code.
mkdir -p run && cp report.json diff.patch run/
python3 tryout_plate.py run
echo "exit=$?"
A zero means the plate accepted this tryout. It does not mean the function is correct for every integer. It does not mean a reviewer can skip the diff.
The checker never executes the patch. Students still run the unit test themselves and write tests_exit from that process, not from the agent's claim. The second command is the gauge the first command only reads.
python3 -m unittest test_clamp.py
echo "tests_exit=$?"
An agent that announces success without a process exit code fails the plate on tests_not_executed. Speech is not a gauge. That single gate is the workshop's real lesson.
A false pass has a smell the room learns to name. The checker printed pass, unittest exited zero, and the diff still looks wrong because max(0, ...) changed a legal span that was already inside the limit. The plate did not lie. The fixture was too thin.
Students add one assertion, self.assertEqual(clamp_span(2, 5, 10), 3), rerun unittest, and only then copy the new exit code into report.json. The repair is a better plate, not a softer gate. A hosted model is not invited to edit the test in order to make the smell go away.
The ninety minutes move in five beats. The first ten belong to the scene: a green badge is not a merge, and the live documentation, not a blog number, sets the free-tier bound for the day. The next fifteen belong to the fixture.
Students run the unit test once, then confirm the checker fails closed when report.json is missing. Failure closed is the behavior they want before any model is involved. A missing file that exits zero would train the room to trust silence.
From minute twenty-five to minute fifty, each pair sends the fixed task through free model access. The prompt stays dull. Fix clamp_span so a negative span raises, keep the edit inside clamp_span.py, stop when unittest exits zero, and write the stop reason into report.json.
Pairs may not improve the prompt. A tryout that changes the ask is a different die. Pairs on small laptops use the free server option for that same prompt and download the transcript when the run stops.
The next twenty minutes are scoring. Each pair runs the checker, then breaks one gate on purpose: a second file in files_touched, a true tests_ran with no local unittest, or a patch that contains a made-up key. Three failures, three reasons, then one pass restored by reverting the break.
The last twenty minutes name what a hosted log can hide, and what this plate cannot see. A logic bug that still exits zero walks straight through. The plate catches process lies and boundary violations.
It does not catch wrong but plausible arithmetic. A human still reads the diff before anything leaves the lab branch. The free server made the attempt easy to repeat. It did not make the reading optional.
Two exercises are enough for a later rerun. Exercise A takes fifteen minutes. Delete stop_reason, confirm the verdict flips to fail, then restore the field.
Exercise B takes twenty minutes. Point the same checker at a second run directory from the free server, without editing the gates. If the export shape differs, students write a short adapter that emits report.json.
Adapters may change shape. Gates stay boring. A third, optional pass takes ten minutes the next morning: run the checker on an empty directory and confirm it raises rather than printing pass. Fail closed, even when the room is in a hurry.
This approach is a lab plate, not a release process. It should not be used on repositories that hold credentials, customer data, or unpublished vulnerability detail. A free server is still a remote machine, and this article does not claim a retention policy, an isolation guarantee, or a permanent quota. Read the current terms before any file leaves the laptop.
Teams that need a named model build, an attested log, or a fully offline room should not depend on a free hosted option for the run. They can still reuse the checker. Instructors who want a leaderboard will be disappointed.
The script records pass or fail for one fixture. It does not measure tokens, latency, or model rank, and those numbers are not invented here to fill a chart. A class that needs those measurements should design a separate bench, with clocks they control, after this plate is already boring.
The same plate refuses a certain optimism. Free model access removes the excuse that only one student could afford to try. It does not remove the need for a stop condition that a second process can see.
An earlier workshop on this account taught an agent loop how to stop. This one teaches the room to notice when the stop was only a sentence. The die is proved on the plate, or it is not proved.
Readers who want that split — a hosted attempt, a local grade — can check MonkeyCode's current free model access and free server option, copy the live limits onto a whiteboard, and rerun tryout_plate.py against a directory they produced themselves. The plate stays useful if another runner is swapped in. That is the property a toolroom actually keeps.