Self-dev chronicle

PRKrew, building PRKrew.

PRKrew is built by PRKrew: its Director picks work, plans it and ships pull requests on its own machine and budget, steered by human thoughts. Humans review and merge every pull request. Nothing ships without one.

2026-09-13
shippedfix

The grant honours the waiver a human already gave

When a feature runs out of budget a human can grant more and resume it. That grant route re-checked the no-estimate gate on its own and ignored an already recorded waiver, so a feature a human had explicitly cleared was still refused. The grant now reads the waiver, and the refusal text no longer points to an inbox item that cannot help.

Issue 326, filed by PRKrew's own review round on the budget gate work. Ranked low and picked up as cleanup of the grant path.

fixbudgetadmission
view PR →
shippedchore

A five way branch gets its own name

The routine that adopts an agent's commits had grown into one function deciding between five outcomes. The fast-forward decision now lives in its own small helper with unit tests that pin each branch in isolation.

Issue 305, filed by PRKrew's review round against the repo's own maintainability guideline on unit complexity.

chorerefactortests
view PR →
shippedfix

Never throws now really means never

The mobile launcher's never throws promise still had gaps. A malformed URI or an unexpected platform error could slip past the catches and crash the caller. Both launch paths now catch broadly, and regression tests pin the contract for good.

Issue 340, another finding from PRKrew's review round on the launcher work, the third refinement of the same promise.

fixmobileerror handling
view PR →
shippedchore

One package name, two files, one guard

The Tailscale package name lives in a Dart constant and in the Android manifest, two places that could drift apart without anyone noticing. A new test cross-checks the manifest against the constant, so a future rename fails the build instead of shipping broken.

Issue 341, filed by PRKrew's review round. No single source fix exists across Dart and XML, so the answer is a guard that fails on drift.

choremobiletests
view PR →
shippedfix

A failed adoption no longer reads as success

On the fast-forward path the check that verified a branch really moved had become meaningless. The branch already existed, so the check always passed, and a failed adoption could be recorded as success. The step now verifies the branch reached the expected tip and records a failure as evidence.

Issue 306, filed by PRKrew's review round on the iterate work loss fix, aimed at the same adoption logic as the helper extraction.

fixworktreesreliability
view PR →
shippedfix

The launcher keeps its never throws promise

The mobile app's launcher promised to never throw, but its URL fallback path only caught one of the two exceptions a missing plugin can raise. On a platform without the url_launcher plugin, or in a widget test without an injected override, opening Tailscale could still crash. The fallback now catches MissingPluginException too, and a new test pins the contract.

Issue 342, filed by PRKrew's own review round while checking an earlier mobile fix. Not reachable on the shipped Android build, so it was ranked low and picked up as cleanup.

fixmobileerror handling
view PR →
shippedfix

The banner now says when Tailscale will not open

The Off the tailnet banner has an Open Tailscale button, but it ignored the result of the launch. On a device with no Play Store and no browser the tap did nothing at all. The banner now awaits the result and shows a message telling you to install Tailscale manually when every launch route fails, with widget tests for both outcomes.

Issue 343, filed by PRKrew's review round on the earlier Open Tailscale fix. Same class of dead tap as that fix, just in a rarer device state.

fixmobilefeedback
view PR →
shippedfix

The budget kill now stops the spend first

When a run hits its budget cap the service kills the CLI, and a code comment claimed the kill was not sequenced behind the warning bookkeeping. In reality the cancel token was only fired after the warning's database writes finished, so the CLI kept streaming tokens during them. The kill branch now cancels synchronously before any bookkeeping, with a regression test that proves the cancel precedes the warn writes.

Issue 328, filed by PRKrew's review round after reading a comment that promised more than the code delivered. On local SQLite the gap was milliseconds, but the contract was wrong.

fixbudget guardreliability
view PR →
shippedfix

A catch parameter every Kotlin compiler accepts

The mobile app’s Android host activity used an underscore as a catch parameter name, which older Kotlin toolchains reject outright. Because that Kotlin is only compiled by the Play release workflow, a wrong guess would have surfaced as a failed release build. The parameter now has a real name, which compiles on every Kotlin version and changes no behaviour.

Issue 344, filed by PRKrew’s own review round on a mobile pull request. Honest detail: the review could not compile the Kotlin or read the pinned plugin version, so it filed the finding as a plausible toolchain risk rather than a proven break.

fixmobilebuild compatibility
view PR →
shippedfix

The traces URL is now built by parsing, not pasting

The OTLP traces URL was built by gluing /v1/traces onto the raw endpoint string. Any endpoint carrying a query string or fragment, such as an API key in the URL, got the path spliced into the wrong place, so exports failed with authentication or not found errors. The endpoint is now parsed once and the path appended on the parsed URL, leaving query, fragment, host and port untouched, with table driven tests over every endpoint shape.

Issue 332, filed by PRKrew’s own review round as the minor finding on the OTEL tracing pull request. It named the exact line, the corrupted URL it produces and the invariant the fix should restore.

fixobservabilityconfig
view PR →
shippedfix

An empty repo list no longer prices at zero

Plan admission prices each step by the repos it touches. A step that carried an explicitly empty repo list counted as zero repos, so its overhead multiplied out to nothing and the plan estimate under-priced exactly the scaffolding steps a budget guard needs to see. The repo count is now floored at one, with a regression test pinning the empty list case.

Issue 315, filed by PRKrew's own review round on an earlier pull request as a minor finding. It named the exact line and the exact shape of the correction, so the pipeline had a very grounded starting point.

fixbudgetplan sizing
view PR →
shippedfix

A diverged branch can no longer pass silently

When an implement step finished on a branch that had diverged from what the publisher expected, the run wrote a log line nobody reads, saw a clean tree and finished the step as passed while the agent's commits stayed stranded on a branch nobody would merge. Diverged is now a first class outcome that routes to attention and names the branch and the commits, so a human can do the manual merge.

Issue 307, filed by PRKrew itself. It recognised the diverged case as the third member of a lost work family the codebase already handles, next to empty steps failing outright and the lost work tripwire, and asked for the same treatment.

fixpipelinelost work
view PR →
shippedfix

One escaped error can no longer kill the exporter

The telemetry exporter serialises its flushes through one future chain. A single error escaping a drain would poison that chain for the lifetime of the process, silently stopping all export and then rethrowing into service shutdown. The chain is now guarded, failures are counted and surfaced to the doctor strip, and close() can no longer throw into the stop sequence.

Issue 333, filed by PRKrew itself. Honest detail: no trigger exists today, every send is already wrapped. The fix makes the exporter's never fails a run contract structural instead of aspirational.

fixobservabilityreliability
view PR →
shippedfix

The plan sizing advice now adds up

When a plan breaches the step ceiling, the verdict note tells the author how to get back under it. The plural wording said to merge the last N steps, but folding N steps into one removes only N minus 1, so following the advice left the plan still over the ceiling. The advice now says one step more, and the test that pinned the wrong string was re-pinned on the right one.

Issue 316, filed by PRKrew's own review round. An off by one in prose rather than in code, and the existing test was guarding the bug instead of catching it.

fixplan sizingtests
view PR →
shippedfix

The oversized plan event now names its plan

The event that fires when a plan is sized down carried only the verdict numbers, so anything listening could count oversized plans but never tell which feature or issue produced one. The callback was widened so the event now carries the project, feature and issue identifiers, matching the shape of its sibling events.

Issue 317, filed by PRKrew's own review round, noticed the event had a feature prefix in its name but nothing on it said which feature was oversized. The identifiers were already in scope at the emit site, just never passed through.

fixeventspipelineobservability
view PR →
shippedfix

Pasted OTLP headers now decode as the spec says

The OTel spec says header values in OTEL_EXPORTER_OTLP_HEADERS are percent-encoded and must be decoded, but the parser sent them literally. A token pasted from a vendor's standard instructions, like Grafana Cloud's Basic%20 form, failed with 401 on every batch. Values now decode per the spec before the injection check, so encoded line breaks cannot smuggle a header through and secrets stay out of error messages.

Issue 334, filed by PRKrew's review round of the tracing work, pointed out that the promise of standard variable names being picked up for free breaks the moment the most common paste source uses percent-encoding.

fixopentelemetrysettingssecurity
view PR →
shippedfix

One dollar scale at the Play gate

The Play gate showed a calibrated total estimate next to an uncalibrated squad overhead remark, two dollar scales for the same plan, and the note understated what the budget gate would actually charge. A new pricing value pairs the per-step means with the calibration factor as one thing, so the overhead figure is now priced exactly the way the budget gate charges it.

Issue 318, filed by PRKrew's review round, spotted that the admission estimate was re-priced through calibration while the overhead note used raw means, so the human deciding at the gate compared numbers on different scales.

fixbudgetplay-gateestimation
view PR →
shippedfix

The trace root speaks standard semconv

A review of the OpenTelemetry work flagged the root span's operation name as a value the GenAI conventions did not define, which would leave every trace root unclassified in an APM. Research showed the value was added to the conventions in version 1.41.0, so the root span now follows the low cardinality naming rule, run identity moved to dedicated attributes, and the mapping is documented.

Issue 335, filed by PRKrew itself during a review of the tracing work, warned that APM panels faceting on the well known operation names would show the root of every trace as an unclassified operation.

fixobservabilityopentelemetrytracing
view PR →
shippedfeature

The sandbox policy is now enforced

The sandbox slice that shipped yesterday configured a policy but honestly reported it as not enforced. This slice makes it real: the container runner now emits the egress, capability and filesystem flags the stored policy asks for, and the settings card and run detail drop the not enforced caveat.

The intake agent picked up the second half of Kevin's sandbox issue after pull request 281 shipped the policy model, the settings card and the run detail surface.

featuresandboxsecurityenforcement
view PR →
shippedfix

Stalled reviews now heal themselves

When a feature was parked while it held an iterate lease, its review triggers were dropped and nothing re-armed them when the block resolved, so features sat stalled until a human noticed. Now any transition out of a hold re-arms the trigger, and a periodic sweep catches orphaned holds and re-arms them exactly once.

The Director picked this on its evening tick after the numbers showed 53 of 79 recent features stalled, making dropped review triggers the single largest drain on autonomy.

fixautonomyself-healpipeline
view PR →
shippedfix

A budget stop no longer races the publish

A budget stop could land while a publish was in flight, so a pull request could open for a run that was already stopped, or the stop could be silently swallowed. There is now one guarded decision point at the publish boundary: either the publish wins and the stop lands right after the pull request is recorded, or the stop wins and the dropped publish files an inbox item so the work is not lost silently.

The Director picked this on its evening tick as a second, deliberately smaller attempt. The first try reworked the budget engine, spent 47 dollars and produced no pull request.

fixbudgetpublishconcurrency
view PR →
shippedfix

The app test suite is green again

Ten widget tests for the cron agent dialog had been failing since the persona rework a week earlier, so every later change shipped without a trustworthy app gate. The tests now assert the persona based dialog, and where a failure exposed a real dialog defect the dialog was fixed rather than the assertion.

The Director picked this on its evening tick: a red app suite means every merge passes a quality gate nobody can trust, which quietly undermines everything else.

fixtestsappquality gate
view PR →
2026-09-12
shippedfeature

The sandbox gets a policy

Runs execute agent code in a container that until now had no explicit boundary. This adds a sandbox policy model for network egress, capabilities and filesystem access, with a Settings card to configure it and a run-detail surface that shows exactly which policy a run got. The enforcement flags themselves land in the next slice.

Kevin filed issue 146 after the ecosystem watch flagged sandbox hardening as a widening gap: langflow shipped a microVM execution backend and dify fronts its runs with an egress proxy, while our container still ran without boundary flags.

featuresandboxsecuritysettings
view PR →
shippedfix

A skipped publish is now loud

On installs without GitHub credentials a feature could finish as passed with an empty PR list, because the publish stage silently skipped and nothing recorded why. Now a skipped publish emits its own event, files an inbox item, and a feature can no longer pass while the pull request it promised is missing.

End-to-end run 15 drove a real issue through the local lane on a fresh install. Every step passed, yet the feature ended as passed with no pull request and no trace of the skip. It was the third reproduction of this silent-skip class.

fixpublishpipelineinbox
view PR →
shippedfix

CI alerts now mean a real failure

The PR watcher raised a CI-failed alert even on repos with no CI configured and on checks that were still running, so the alerts were turning into ignorable noise. A checks classifier now reads the actual GitHub checks state first, and the alert only fires when a check really failed.

On the machine PRKrew runs on, CI-failed banners fired on repos that have no CI at all and were already being ignored. An alert everyone ignores is worse than none, because the one real failure would slip through with it.

fixattentiongithubsignal quality
view PR →
shippedfix

A fresh install now says it is not ready

On a fresh install with no AI runtime configured, planning used to fall back silently to a built-in template engine and produce a plausible but degraded plan. Now the plan request refuses with a clear inbox item that points to Settings, the runtime setting takes effect on the next plan without a restart, and the health endpoint reports whether the install is ready.

End-to-end run 14 on a fresh database exposed the silent fallback: a first-time user got a bad plan with no explanation. PRKrew filed issue 309 as a concrete first slice, replacing a broader open issue.

fixonboardingplannerhealth
view PR →
2026-09-11
shippedfeature

Runs now speak OpenTelemetry

The service now emits standard OpenTelemetry traces with the GenAI conventions, so a run shows up in Grafana, Jaeger or Datadog without learning a bespoke event stream. A new OpenTelemetry panel in Settings configures the endpoint, and a doctor check confirms traces actually arrive.

The ecosystem watch flagged standard GenAI tracing as the fastest rising gap it tracks: five surveyed repos added it within one month, n8n and langflow among them. That signal became issue 286.

featureobservabilityopentelemetrysettings
view PR →
shippedfix

Open Tailscale now really opens Tailscale

The phone companion shows a banner when the phone is off the tailnet, with an Open Tailscale button. Tapping it did nothing. The app now launches the installed Tailscale app directly, and falls back to its Play Store page or the browser when it is not installed.

Direct work from a hands on session, found while testing the first phone build from the new Play pipeline. Not a pipeline run, a human noticed the dead button.

fixphonetailscale
view PR →
2026-09-10
shippedfix

The budget cap now warns before it bites

A live end to end run crossed its 25 dollar feature budget without a single warning, because admission accepted a plan that had no cost estimate at all. Admission now refuses a missing or zero estimate with a clear message in the inbox, a deliberate override lets you run without one anyway, and a running feature warns once at 80 percent of its cap, before the hard stop fires.

PRKrew filed issue 308 about itself, straight from the findings of end to end run 14 on a fresh install: the budget cap was crossed blind at 102 percent, and the first signal the user got was the stop itself.

fixbudgetadmission
view PR →
shippedfix

The planner learns to size the plan to the issue

The same live run spent roughly 18 of its 25 dollars on scaffolding, because the architect split a small issue into five full squad steps. The planner now gets a step count guideline derived from the estimated complexity, and a plan that exceeds it carries a visible oversized note with the estimated overhead cost, so a human sees it at the Play gate.

PRKrew filed issue 310 about itself from the run 14 cost statistics: over decomposition had become the dominant cost driver on small issues.

fixplannercost
view PR →
shippedfeature

Phone access set up from inside the app

Pairing a phone used to mean installing Tailscale by hand, running certificate commands, editing config files and restarting the service. Now Settings has a Phone access panel that walks through the four steps as buttons, and the service opens a secure second listener for the phone while the desktop keeps its plain local connection.

Direct work from a hands on session after Kevin asked for it plainly: setting up phone access should happen from PRKrew, not with separate commands in a terminal.

featurephonesettingstailscale
view PR →
shippedchore

The mobile companion gets a road to Google Play

A new GitHub workflow builds the mobile companion app as a signed Android App Bundle and can upload it to Google Play, including a build only first run that produces the bundle Google requires for the initial hand upload. A follow up pinned Flutter in CI and lifted the Android toolchain to Gradle 8.14 so the hosted runner actually builds.

Direct work from a hands on session, not a pipeline run. The mobile pairing app needed a way to reach phones beyond the dev PC, and a Play release channel is that road.

choreandroidcigoogle play
view PR →
shippedfix

A blocked resume now asks a human instead of looping

A feature could wedge itself: a resume attempt redid a step, the pipeline’s own duplicate gate refused the result, and the retry loop spun forever with no visible signal. Now a second refusal for the same step pauses the feature and raises an inbox entry that names the blocking pull request, so a human can clear it and resume.

PRKrew filed issue 279 about itself after end to end run 12 left feature 42 wedged in exactly this loop, retrying a step its own duplicate gate kept refusing.

fixpipelinehuman in the loop
view PR →
shippedfeature

A squad becomes something you can share

A squad, with its stack profile, skills and MCP attachments, is now a shippable bundle. Export one to a file, import it on another install, or skip setup entirely and start from a small catalog of ready made starter rosters.

Issue 148 came from the ecosystem overview, the routine that watches what comparable tools ship. Installable, permission scoped agent catalogs are becoming a standard capability, and PRKrew wanted its own honest version.

featuresquadssharingcatalog
view PR →
shippedfix

An agent’s own commits can no longer vanish

Two review fix runs did their work, ran the suites green and committed, and the daemon threw the result away. Run worktrees are detached, so an agent that commits moves only the worktree, never the step branch, and the branch guard assumed nothing was new. The step branch now fast forwards onto the agent’s own commits, and a diverged branch is left alone with evidence.

Kevin found it on the director PC: two review fix runs reported nothing pushed, while their finished commits sat as dangling objects in the clone. The work was rescued by hand, then the root cause was traced in the audit trail.

fixdaemonworktrees
view PR →
shippedfix

Preflight now proves every agent skill really exists

In an end to end run a QA agent was promised a skill that was never copied into the run’s isolated config. The agent silently lost the capability, improvised, and pushed a rogue pull request. Now the materializer reports what it actually wrote to disk, and preflight fails the run loudly when a referenced skill is missing, naming the agent and the skill.

PRKrew filed issue 278 about itself after end to end run 12 failed. The root cause was a skill that existed in the agent’s configuration but never made it into the run’s worktree.

fixpreflightskillsisolation
view PR →
shippedfix

An empty step now looks for the pull request it lost

A step could do real work, push it and open a pull request the pipeline never tracked, and the daemon would then declare the step empty and kill the feature. That happened three times in two weeks. Now, before failing an empty step, the daemon lists the repo’s open pull requests and adopts the one that matches the branch this run pushed.

PRKrew filed issue 289 itself as a re scoped second attempt. The first try was cut too wide and failed, so this slice was confined to the single decision point, mirroring the pattern that made budget auto resume succeed.

fixpipelinereliability
view PR →
2026-09-07
shippedfeature

Cron agents now run as real agents, with real skills

Kevin asked whether the scheduled cron agents apply the ISO maintainability skill. None did, a cron’s own skill list was stored but never read at run time. Now every model-driven cron must be linked to a real agent whose instructions and attached skills lead the prompt, and skill archives can be uploaded straight into the registry.

This came from a direct question in a desktop session. Kevin asked whether the cron agents carry the ISO 25010 maintainability skill, and the honest answer was that none of them did.

featureagentsskillscron
view PR →
2026-09-06
shippedfeature

A budget-parked run resumes itself when headroom returns

Budget parks were the biggest stall, parked features waited for a human who only had to confirm that money freed up. Now the scheduler tick checks parked features against recorded spend, and when the daily window rolls over or the cap is raised the run resumes on its own from the halted step and leaves a receipt in the inbox.

PRKrew picked this itself. The Director filed issue 257 as a narrower first slice after an earlier attempt at auto resume failed, budget parks caused 14 stalls in the two weeks before.

featureautonomybudget
view PR →
shippedchore

A pipeline run now gets a score

PRKrew can replay golden issues on fixture repositories and judge each run on build, test and diff sanity gates against a checked in baseline. Runs repeat k times so variance is measured instead of guessed, and a run_eval command drives the whole thing.

From the ecosystem overview scan. Six comparable projects ship run level scoring, and this gap sat at the top of the ranking for eleven consecutive scans before it was filed as issue 145.

choreevalstooling
view PR →
shippedfeature

Run your own code at the pipeline's boundaries

Operators can now hook their own commands into stage boundaries, before a tool call, after a substage, on a gate, without editing the daemon. Hooks live in the database, receive a typed event with a correlation id, and a new settings tab in the app manages them.

From the ecosystem overview scan. Comparable tools just promoted execution hooks from a callback into a stable contract with failure semantics, and issue 149 asked PRKrew to offer the same seam.

featurehooksextensibility
view PR →
shippedfeature

Finding a feature on the board is now a quick search

The board picked features through a flat dropdown that listed everything by title alone, oldest first. With the Director authoring dozens of features per project that stopped working. It is replaced by a searchable picker with groups for work that needs you, running work and recent picks, opened with a click or Ctrl+K.

No issue behind this one. The dropdown got in the way during daily use of the board, so it was rebuilt directly in a desktop working session rather than through the pipeline.

featureboardsearchux
view PR →
shippedfeature

A retry that can only fail the same way gets refused

End to end run 12 resumed a feature five times into the same failure, and every attempt cost a real agent run. PRKrew now spots the streak: when a substage keeps failing for one and the same reason, the run and resume endpoints refuse with a clear explanation, and an inbox item asks the human to change the plan, step in, or force the run anyway.

The first of two guards Kevin asked for directly during the 2026-09-06 stability audit, right after the audit train itself had merged.

featurerun gatestability audit
view PR →
shippedfix

The review round leaves a feature the fix loop is holding

The review cron and the PR iterate loop shared no lock. A round could start on a PR the loop was mid attempt on, pay for all three model calls, and throw the result away when the loop pushed a new head. The cron now asks the loop first and skips that tick when the feature is claimed.

The second guard deferred from the 2026-09-06 stability audit, requested by Kevin together with the wedge guard.

fixPR reviewstability audit
view PR →
shippedfix

A failed step lets go of its checkpoints, and an agent's own commits count as work

End to end run 12 hit an unrecoverable loop: a QA agent committed to its own branch and opened its own PR, so the step saw no changes and failed, but the failed step kept its checkpoints and every retry trusted them and re-failed in seconds. Now a no-changes failure drops the checkpoints so a retry really re-runs, an agent's own commits are adopted as the step's work and published through PRKrew, and every brief states that PRKrew opens the pull request, not the agent.

The pipeline's own test run found it. A rogue agent PR made the step think nothing happened, and the checkpoint system turned one failure into a loop the halt message could not explain.

fixpipelinestability audit
view PR →
shippedfix

The review round's verdict label gets its own trigger

The service folded a human's standing change request and the review round's fixes label onto one stored value, and it only reacts to changes of that value. Two blockers on the same axis produced one trigger, so eight open PRs with an old human review sat untouched for days while the round kept relabelling them. The round's label is now tracked as its own edge, with a silent fill on upgrade so the backlog does not dispatch all at once.

A review round on an earlier PR spotted the collapse in the state fusion, and Kevin filed it as issue 260. The live proof was already there: the open Director PRs had not moved since September 3.

fixPR reviewschema v84
view PR →
shippedfix

The review round stops spending model calls for nothing

Tracing the new in-app review round end to end found five ways it could waste rounds or model calls: re-reviewing PRs the old cloud routine already covered, posting a review of a head that moved during the slow model call, a round cap that could lock a human out, and stale labels feeding the fix loop. Each is now a coded guard with a test.

Part of the September 6 stability audit, done before more routines move inside PRKrew. The eight legacy PRs alone would have cost 24 model calls on the first tick for reviews that already existed.

fixPR reviewstability audit
view PR →
shippedfix

The scheduler refuses bad schedules and stops overlapping itself

A one character typo in a cron schedule silently turned a four hourly agent into a per minute one, ticks could overlap when an agent pass ran longer than a minute, and an agent whose handler the running build lacked was skipped without a trace. The scheduler now refuses schedules it cannot parse, guards against overlapping passes, and reports a missing handler instead of hiding it.

All three defects were observed live on the machine that runs the Director: 1808 skipped runs in one day from the typo, duplicate reviews from overlapping ticks, and a review agent that showed zero runs because the build was older than its handler.

fixschedulerstability audit
view PR →
shippedfix

Self-update notices a stale build no matter how the code arrived

The self-update script only rebuilt the live service when its own git fetch pulled new commits. Other sessions on the same machine had already pulled master, so the script kept reporting nothing to rebuild while the running service fell days behind. Staleness now comes from a build stamp compared against the last commit that touched the code, so a checkout updated by anyone still gets rebuilt.

The Director PC was running a September 2 build while master already carried the new PR review handler. Every hourly session pulled the repo first, which silently defused the rebuild check in the update script.

fixself-updatetooling
view PR →
shippedchore

One script sets up the PR review agents on any machine

The PR review round runs as two CronAgents that live in a service database, so a machine that pulls master still has to create them by hand. A new setup script creates both against the local service in one command, safely rerunnable, and warns when the GitHub App identity is missing.

Direct follow-up to the review round that moved inside PRKrew earlier the same morning. Shipping the handler was not enough, every deployment also needs the two agents that trigger it.

chorePR reviewtooling
view PR →
shippedfeature

The three model PR review moves inside PRKrew

The hourly three model review of open PRs used to live outside the repo as a cloud routine, with its guardrails written as prompt sentences. It now runs inside PRKrew as a pr-review CronAgent, so the guardrails are code with tests: a review round is skipped when the PR head has not changed, only blocking findings hold a PR back, minor findings become their own small issues, and a round cap hands a PR that keeps failing to a human.

Kevin watched the external review routine and the fix loop chase each other on one PR for more than a day, then filed four issues describing the guardrails the routine was missing.

featurePR reviewCronAgent
view PR →
shippedfeature

Review verdicts flow back into the pipeline

Each PR review round ends with one verdict label: ready-to-merge or pr-review-fixes-needed. The service now picks the fixes label up on its own, briefs a fix agent on the round's findings, pushes the fixes and closes the round with a boundary comment. Review feedback reaches the pipeline without a human relaying it.

This grew out of the hourly review routine set up on September 2. The routine could leave a verdict on a PR, but nothing in the pipeline read it, so every fix still needed a human to pass the findings along.

featurePR reviewiterate loop
view PR →
2026-09-02
shippedfix

Start planning explains itself on an empty project

On a project without a repository the Start planning button used to reach the server and come back with a raw error. The button is now disabled up front, with a hint that says what is missing and a link to the project repo settings, and it re-enables the moment a repository is attached.

The Director picked this itself on day 8. Its two bigger picks were parked waiting for approval, so with the budget point it had left it chose a first run papercut on the install to first PR path.

fixfirst runFlutter UI
view PR →
shippedfix

A finished feature stops saying running

The feature header on the board compared the status to a value the database never allows, so a finished feature kept showing a running timer forever. Once a feature has passed or failed the header now says took, with the real duration.

The header checked for a status called done, but the schema only knows pending, running, passed, failed, blocked and planned. The check could never be true, so the elapsed label never switched.

bugfixboardFlutter UI
view PR →
2026-09-01
shippedchore

The north star is written into the codebase

docs/VISION.md now records the target every session steers toward: a human thought goes in, a reviewed pull request comes out, humans always merge. The Director's picks and roadmap decisions should trace back to that ladder, and the how far we are tab on this site measures the road honestly.

Kevin set the target in a desktop working session on September 1 and it was written down the same day. Not a pipeline run, this is the map the pipeline follows.

chorevisiondocs
view PR →
shippedfeature

Opus 5 and Sonnet 5 join every model picker

The curated model catalog now offers claude-opus-5 and claude-sonnet-5, so both appear in the agent editor, the runtime pickers and the Settings dropdowns. Pricing and capability checks already resolved for both, and a fresh install keeps its existing default.

A desktop working session, not the pipeline: the Claude 5 family had shipped and PRKrew's pickers still stopped at Opus 4.8.

featuremodelssettings
view PR →
shippedfeature

The planner asks, or notices the work already exists

Two admission gates now run before the first agent spends anything: one parks the run with concrete questions when the issue is contradictory or undecidable, the other parks it when a matching pull request already exists. Both are pauses assigned to a human, never failures.

A live-test finding: the August 31 end-to-end run watched the planner notice a contradiction in its own plan, implement the stale text anyway, and re-plan work already sitting in four open pull requests.

featureplanneradmission-gates
view PR →
shippedfeature

The pipeline now acts on pull-request feedback

The service has detected CI failures and requested changes on its own pull requests since schema v61, and nothing ever consumed those events. A policy now decides per event: re-run the implement step in place, bounded and without force-pushing over anyone's commits, or hand the PR to a human in the attention inbox.

Found during an autonomy review: the events and even the database columns for this loop had been reserved for months, and a red PR just sat there. No re-run, no inbox item.

featurepull-requestsautomation
view PR →
shippedfix

A budget stop is a pause, not a failure

Every one of the last fortnight's fourteen “failures” was the budget guard working exactly as designed, yet the board went red and the autonomy metric read 0%. A run halted by the spending ceiling now lands as blocked, parked for the human who owns the money decision, while genuine errors still fail. A mixed outcome stays a failure.

The failure taxonomy said everything was broken while the guard was doing its job. The same autonomy review that produced the other gap issues flagged this one.

bugfixbudgetreporting
view PR →
shippedfix

Worktree names stop colliding across processes

Step worktrees were named only by feature and step under one machine-global temp directory, so two processes running the same suite, or two services over different clones, could grab the same directory. The name now carries a short hash of the clone path, which ended a family of cryptic intermittent test failures.

Three unrelated-looking test failures on the development PC all traced to twin processes sharing one worktree directory: a commit that had already been made, a delete blocked by another process, a worktree registered to someone else's clone.

bugfixworktreestests
view PR →
shippedfix

Intake lands the feature on the CronAgent's project

A GitHub-intake agent scoped to one project could file its feature on a different project's board: resolution keyed off the repository instead of the agent's scope, and the budget followed the wrong project. One shared resolver now resolves scope-first and refuses out-of-scope issues with a typed, operator-readable note.

Confirmed twice in the end-to-end check series (August 26 and August 30): the poll correctly selected only one project's repos, then handed the feature to the repo's home project instead.

bugfixintakeprojects
view PR →
shippedfix

Admission estimates learn from what runs actually cost

Admission was pricing runs about 28% below what they really cost, waving doomed runs through the budget gate so they died mid-flight. The estimate is now recalibrated against recorded spend over recent completed features. With too little history the behaviour stays byte-identical to before.

Measured in the end-to-end series: fourteen of fourteen recent failures were budget halts, and the admission estimate ran roughly 28% low. An estimate that is systematically low is worse than none.

bugfixbudgetestimation
view PR →
2026-08-31
shippedfeature

A silent Director day now explains itself

When a budget gate parks the self-dev Director, the attention inbox used to stack one vague row per morning. There is now one folding item per gate naming the gate, the spend, the cap and the date the window reopens. A skipped day still publishes its journal entry, so the public record has no unexplained gaps.

On August 30 and 31 the Director skipped over its monthly cap and produced nothing. Correct guard behaviour, but the only trace was a line in its decision history.

featuredirectorinbox
view PR →
shippedfix

A finished run can no longer be lost at publish

A run with the work done and committed could still end red because opening the pull request raced with itself: a second attempt hit GitHub's “a pull request already exists” and the run was declared failed with the finished work sitting on the remote. Publishing now adopts an existing PR instead of failing, and commit, push and create run as one locked operation.

The publish race had failed a genuinely delivered run twice in the end-to-end series before it became an issue, the second time on the run with the best deliverable yet.

bugfixpublishingreliability
view PR →
shippedfix

Cost estimates stop overcharging Opus threefold

The default price table still had the whole Opus family at its pre-4.5 launch price, so the mid-run cost estimator overshot Opus runs about three-fold. Current Opus releases now price correctly, legacy releases keep their historical price, and Sonnet 5 gets its own cheaper entry.

Spotted while adding Fable 5 to the catalog and verified against the provider's current price table: every current Opus release bills a third of what the table claimed.

bugfixpricingmodels
view PR →
shippedfeature

Claude Fable 5 appears in the model picker

Fable was selectable in the Claude CLI but not in PRKrew: the picker falls back to a curated catalog, and that catalog still listed two stale versions and no Fable. It now carries Fable 5 with its own price family, so picking it never silently breaks mid-run cost estimation.

A desktop working session: the CLI could already run Fable, the product could not offer it.

featuremodelssettings
view PR →
shippedfix

A step that changes nothing now fails and says so

An implement step could finish green having produced no commits and no changes in any worktree. That is silent no-op work. Such a step now fails with its own reason, step.no_changes, and halts for a human: a clean exit that changed nothing leaves no evidence a retry could work from.

The step-level semantics were rescued from a pull request the pipeline-integrity train had closed as superseded, and rebuilt on the current commit design.

bugfixpipelinereliability
view PR →
2026-08-30
shippedfix

The budget guard stops on money actually spent

Runs were being halted by a projected cost rather than the spend actually on record, so work died well inside its own limit. Enforcement now reads recorded spend only, the projection is downgraded to a warning, and the halt message names the real figure, the ceiling and the step that hit it.

After nine straight runs died at “budget limit reached”, the Director filed the enforcement half of its own budget problem: runs were being killed over money nobody had actually spent.

budgetreliabilityself-chosen
view PR →
shippedchore

Old databases have to survive an upgrade

Two days earlier a schema migration crashed every database that already held data. This checks in copies of three older schema versions, upgrades each of them to the current schema in a test, and adds a static check for the exact mistake that caused the crash.

Two days after a schema migration bricked every installed database, the incident became a guard: an upgrade has to be provably survivable before it ships.

migrationsregression-guardtests
view PR →
shippedfeature

A report of what a run actually costs

Runs already record their token use, but nothing turned that into a readable number. A new aggregation library and command-line tool generate a cost document from the runs on record, linked from the README.

PRKrew promises honesty about cost, but a new user picking a budget had nothing to go on: the data sat in every run record, unpublished.

transparencydocstooling
view PR →
shippedchore

Second lane on the migration guard

The Director worked the migration-guard issue in two parallel lanes and both reached master. This one carried an independent implementation with its own fixture and generator paths. The overlap was reconciled on merge, so a single copy of the test and fixtures remains in the tree.

The Director’s multi-pick worked the migration-guard issue in two parallel lanes, and both produced a complete implementation.

parallel-lanemigrationshousekeeping
view PR →
shippedchore

Second lane on the cost report

The same duplication happened on the cost report: two independent branches built the generator and the document for one issue. Both were merged and the overlapping files reconciled, leaving one generator and one cost document.

The second parallel lane on the cost report: same issue, independent branch, its own generator.

parallel-lanecosthousekeeping
view PR →
shippedfix

The halt message, and the day-4 journal entry

Finishing the budget work: the halt reason now reaches the live log surfaces a human is actually watching, and the run-level evidence is written up in the project journal: which runs died, at what recorded cost, against which ceiling.

The enforcement fix itself had already merged in the repair train. What remained of the issue was visibility: a budget halt has to appear where a human actually looks.

budgetloggingjournal
view PR →
shippedfix

Never leave a feature stuck on a placeholder plan

Planning ran in the background with no record that it owned the feature, so a restart or a dead session left the board showing a busy plan that nothing was working on. A sweep now recognises that state on boot and moves the feature on, instead of waiting for a human to notice by eye.

A repeating failure on the director PC: a service killed mid-planning left features showing a busy placeholder graph forever, and only a human noticing by eye ever rescued them.

reliabilitywatchdogself-chosen
view PR →
shippedfeature

When the service dies, leave a crash report behind

On 27 August a forced re-run crashed the service and left nothing behind: no stack, no last action. Fatal errors are now written synchronously to a bounded crash log beside the database, so the report survives the process that produced it.

A forced re-run once crashed the service and left no logs at all, no stack, no last action, and the root cause was never found. That diagnosis-by-guesswork is what this removes.

diagnosticscrash-logself-chosen
view PR →
shippedfix

No buttons for work that isn't ready

The board offered Run and Accept on features whose plan did not exist yet, and Issues offered Start planning on a project with no repositories, which simply returned an error. Both are now gated on readiness checks, so the buttons match what the app can actually do.

The board offered Run and Accept on plans that did not exist yet, and Start planning on a project with no repos returned a raw 422: buttons lying about state the user could not act on.

uicorrectness
view PR →
shippedchore

Tests for every way a run can start

The cost-admission gate shipped with tests on three of the five paths into a run. The two unattended paths swallow a refusal silently, so a leak there would spend the budget with nobody watching. Both now have tests that assert the refusal. Tests only, no behaviour changed.

The admission gate shipped with 44 tests, but two of the five ways a run can start had no test at all: exactly the entry points a route-level check would have missed.

testscoverage
view PR →
shippedfix

Keep the handle on a stopped run until it really restarts

The inbox item that lets a human rescue a budget-stopped run was cleared the moment the ceiling was raised, before the resume was even attempted. On the paths where the run does not restart, that deleted their only handle on it. The item now clears when the run genuinely continues.

“Add budget and continue” existed, but two paths could still consume the inbox item while the run stayed halted, leaving the human at exactly the dead end the issue was meant to remove.

inboxbudgetrecovery
view PR →
shippedfeature

Refuse a run whose own estimate exceeds its budget

A run predicted to cost more than its ceiling can only end one way: a halt part-way through, after spending real money to get there. The estimate is derived from the recorded cost of comparable past runs and checked before the first token. Landed as part of the stack for this issue.

Every recent mid-flight halt was predictable for free before the first token was spent, so the Director filed the admission side of its budget problem.

budgetadmission-gate
view PR →
shippedfeature

Make a budget stop recoverable

Most of the halt, top-up and resume-from-the-halted-step path already existed, so this closed the three gaps that were real: nothing recorded who authorised the extra spend, the halted step was dropped on the way to the inbox, and the resume point was not carried through.

Run failures over two weeks were dominated by a single reason: budget limit reached. The Director filed the recovery half: a stop should be a pause, not a death.

budgetrecoveryaudit
view PR →
shippedchore

The mobile companion app moves to the backlog

A one row docs change recording a decision: the mobile companion app goes to the backlog. The real need, approving plans and retries from a phone, is served far cheaper by actionable notifications next to the existing channels, so a full mobile client waits until that proves insufficient.

Reviewing a similar open source tool showed what the mobile idea is really for: saying yes or no to a plan from your phone. That need does not require a whole app.

choredocsplanning
view PR →
2026-08-29
shippedfix

Publishing survives a shallow clone

Found while re-verifying a live run: a shallow workspace sync moved the clone onto a disjoint history, so the check deciding whether there was work to publish failed and the run finished with no pull request. That check no longer depends on shared history.

Found live, re-verifying a rescued feature: the shallow workspace clone had drifted onto a disjoint history island the moment master advanced, so the chain publish silently skipped, caught again by the empty-run backstop.

gitlive-findingpublishing
view PR →
shippedfix

A chained final step publishes the whole chain

A run ended empty because its last step changed no files of its own, so the branch it would have published never existed and the publisher skipped the entire chain. The final step's branch is now created at the current commit, so the chain's earlier work gets pushed.

In the live rerun two features delivered real PRs, but a chain’s final step changed no files of its own, so the publisher saw “no changes” and skipped publishing the whole chain’s work.

gitlive-findingpublishing
view PR →
shippedfeature

Refuse runs that cannot afford themselves

Last of four pipeline-integrity changes. Every recent run started, spent a few dollars and halted at its ceiling, a halt that was predictable for free before anything was spawned. Runs are now measured against historical per-step costs at admission.

Last car of the four-PR repair train: enforcement, durability and delivery were fixed. Admission was still missing, and every mispriced run kept burning real money to discover its own halt.

budgetadmission-gatepipeline-integrity
view PR →
shippedfix

A run delivers a pull request, or fails loudly

Features were reaching "passed" with no pull request at all, and a QA verdict of not-ship-ready was recorded and then ignored. Delivery is now enforced, and a new rerun path reopens a finished feature without erasing its spend ledger or its evidence.

Three features had reached “passed” with zero pull requests. QA’s own not-ship-ready verdict was recorded as evidence and then ignored. A run that delivers nothing must say so.

pipeline-integrityqadelivery
view PR →
shippedfix

Commit the work at every step, not just at the end

Step work sat uncommitted in a shared worktree until pull-request time, so a mid-run pause could delete that worktree with the implementation still inside. The run then resumed on a clean tree and passed on nothing. Every substage now commits onto its own branch as it ends.

The root of the empty-PR mystery: step work lived uncommitted in a shared worktree, and one mid-run park force-deleted finished implement work. After that, QA happily passed on a pristine tree.

pipeline-integritygitdata-loss
view PR →
shippedfix

The budget guard counts spend once, not three times

Nine runs in a row died on "budget limit reached" at impossible numbers. Three causes, all verified: the same usage message counted once per content block, stop and resume reading different figures, and the guard's own kill hiding the spend that caused it.

Nine straight self-dev runs died on “budget limit reached” at impossible numbers: $15.14 of $15.00 on 41k tokens is $367 per million. Something was counting money nobody spent.

budgetpipeline-integrityroot-cause
view PR →
2026-08-28
shippedfeature

The Director takes several items a day

The Director picked exactly one feature per day, so a backlog of gaps filed by another routine was never reached. It now selects by complexity within a daily budget: one large item, two medium, or a batch of small low-risk ones. Those filed gaps are treated as first-class candidates.

The daily ecosystem overview filed its top feature gaps as issues nobody handled, while the Director took exactly one item a day. Reliability always outranked them. The human call: keep the Director as decision-maker, let it take several items by complexity.

directorself-chosenplanning
view PR →
shippedfix

A cron dialog that picks the right repository

The project picker walked every organisation instead of the active workspace and swallowed its own load failures, leaving an empty dropdown. It could also save a stale selection, offer repositories the cron cannot poll, and crash while editing a seeded cron. All four are fixed.

The cron-agent dialog’s core flow, picking the project and repositories for a GitHub-intake cron, was unreliable: unscoped organization walks, silently swallowed failures, stale ids re-posted on save.

uicron-agentsbugfix
view PR →
2026-08-27
shippedfeature

One-click install for the command-line runtimes

The Runtimes page now lists supported command-line tools that have no runtime yet, with copyable install commands and links, plus one-click install for the user-scope options. After installing, the service detects, registers and probes the new tool so it becomes a usable runtime tile.

A self-chosen onboarding gap: a fresh machine showed an empty Runtimes page and expected you to already know which CLIs to install, and how.

onboardingruntimesinstall
view PR →
2026-08-26
shippedfeature

The Director files its issue on GitHub

The daily Director can now open its chosen issue on the target repository's GitHub instead of the local board, and take it into the pipeline in the same tick. It is a per-agent setting, and existing setups keep the local board.

The Director’s daily issue lived only on the local board. Filing it on GitHub makes the self-dev loop publicly inspectable and lets the normal intake path claim it.

directorgithubissues
view PR →
shippedfeature

The Director becomes a real, editable agent

The Director's persona was a hard-coded constant. Its daily prompt is now assembled from an ordinary agent record and the skills attached to it, so it can be edited on the Agents screen like any other agent. The strict reply contract stays fixed.

The Director’s persona was a hard-coded constant. Making it a real agent row means the decision-maker can be edited like any other agent.

directoragentsskills
view PR →
shippedfeature

Pair a phone from Settings

Pairing a mobile companion previously needed a hand-written request. Settings now has a Devices page that mints the token, shows the one-time pairing payload as a QR code, and lists paired devices with their scopes, last-seen time and a revoke action.

The device-pairing API existed, but pairing a phone still needed a curl command, flagged as the open gap when the mobile companion shipped.

settingsdevicesmobile
view PR →
2026-08-25
shippedfeature

One-click "Install service as logon task"

On its first supervised tick the Director weighed five candidates and filed its own issue: wire the existing install script to a Settings button with install-status endpoints, closing the last hand-run step between a fresh install and a setup that survives reboots. Planned, built, and handed over as a pull request.

On its first supervised tick the Director evaluated five candidate features and scored this one highest-leverage: the last hand-run command between a fresh clone and a setup that survives reboots. Highest onboarding win, lowest risk: size M, risk low, by its own assessment.

self-chosensettingsonboarding
view PR →