2buildDocumentationGitHub
2build / DocumentationRead Markdown (.md) View source

Operations

Day-2 configuration, telemetry, and upgrade handling.

Configuration

bash
bbs config set telemetry local       # off | local
bbs config set update_check true     # false silences upgrade notifications
bbs config set auto_upgrade false    # true runs bbs update on session start
bbs config set proactive true        # false = only run skills typed explicitly
bbs config set foreman_status_interval 3600  # seconds between foreman reconciliation ticks
bbs config set parallel_max_workers 4     # balanced laptop-safe ceiling
bbs config set parallel_global_units auto  # global weighted Foreman capacity
bbs config list                      # show saved YAML (not all effective defaults)

Settings without a dashboard control

These settings are available through the CLI or files; no UI feature flag is needed. Inspect saved user settings with bbs config list or bbs config get <key>. An absent key prints nothing and uses its built-in default; list does not enumerate every default or refresh old header comments.

To open the user configuration directly:

bash
BBS_CONFIG_DIR="${BABYSIT_STATE_DIR:-$HOME/.babysit}"
mkdir -p "$BBS_CONFIG_DIR"
nano "$BBS_CONFIG_DIR/config.yaml"

bbs config set <key> <value> creates/updates that YAML without needing an editor. A key being writable does not mean it is supported: use documented keys below and the configuration examples above.

Setting / fileHow to inspect or changeDefault / scope
semantic_decision_providerbbs config get semantic_decision_provider; bbs config set semantic_decision_provider cloudflarellm; user scope only
semantic_decision_modelbbs config set semantic_decision_model clef-flash or clefclef-flash, used only with cloudflare
Semantic credentialsEdit $BBS_CONFIG_DIR/.env with nano; set CLOUDFLARE_ACCOUNT_ID and CLOUDFLARE_API_TOKENEnvironment wins, then user .env; project .env is ignored
Project inference limitsEdit <repo>/.babysit/semantic-decision.yamlOptional; can restrict external inference, not enable it or select provider/model
Foreman model/effort routesbbs foreman model --json; edit ~/.babysit/settings.json or <repo>/.babysit/settings.jsonBuilt-in policy, then user, then repo overrides; schema
Finish, rigor and repo profilebbs autopilot git-flow; edit <repo>/.babysit/git-flow.yamlProfile policy; separate from semantic provider configuration

The semantic-decision step always exists. Selecting Cloudflare is the explicit user opt-in to external inference; credentials alone do not select it. Switch back with bbs config set semantic_decision_provider llm. No separate semantic_decision_enabled switch is needed. The CLI reads config each call; there is no daemon restart or rebuild for settings changes.

Project restrictions, for example:

yaml
# <repo>/.babysit/semantic-decision.yaml
allowed_kinds: [task-size, task-complexity, testcase]
# enabled: false  # force the LLM provider for this project

Other decision kinds fall back to LLM. allowed_kinds: [] permits no external inference; omitted permits all kinds under the user's chosen provider. Malformed policy also falls back. enabled: true grants no access the user has not configured. Do not run bbs secrets load to source project credentials for this provider; its credentials belong in the user file or environment.

Cloudflare's 20-second timeout and 0.75 confidence threshold are currently code constants, not hidden config keys. Request/response format, all decision kinds and telemetry fields live in the semantic-decision contract.

Decision coverage and human review

Bounded judgments now use one semantic-decision step, with the current LLM as default. This coverage audit separates judgments from enforced policy:

DecisionConsumer / kindWhat still governs execution
Mechanical, Taste, or User ChallengeAuto-Decision Framework / decision-tierKnown User Challenges and missing authority cannot be downgraded
Size, complexity, workflow and decompositionPlan-draft, Foreman routing, Autopilot / task-size, task-complexity, skill-route, orchestrationExplicit user selections, rubric floors, workflow prerequisites
Impacted tests and coverageImplement, QA / testcaseRequired criteria, regression reproducers and repo checks
Finding validity and product severityReview-pr, Fix-pr, Autopilot repair, Foreman product evaluator / review-findingCited disproof before refuting; broken criteria remain material
Extra review worth doingReview-pr / deep-reviewMandatory phases of the selected effort still run
Plan evidence sufficient; human input neededForeman parent/child review, Autopilot / human-reviewArtifact freshness, delegation, holds, grants, safety floors and approval CLI
Recoverable blocker or missing human inputTriage / recoveryRetry budget, ownership and dispatch authority

Foreman's parent plan requires human review unless explicit, persisted --auto delegates it. Child reviews are autonomous by default within the accepted parent plan and current bounds. Semantic judgment evaluates the rubric; approval self-resolve still enforces floor → rubric → authority → approval. A model's proceed answer is not an approval record.

Autopilot honors --stop-after and invoker-held approvals, repairs routine findings itself, and escalates non-derivable User Challenges. Composed Taste choices go in its handoff; it never gains permission to push or release from a semantic answer. Direct skills retain the Auto-Decision Framework's own final reporting gate. The shared human-review contract distinguishes proceed, revise, needs-human, and blocked.

Open-ended design, hypothesis generation and experiment ideation still belong to their domain skills; once they have evidence and bounded alternatives they can use the shared custom kind. Deterministic branch/DAG checks, retry counts, resource admission, artifact hashes and release gates do not need inference.

Machine-global worker admission

parallel_max_workers remains a per-Foreman ceiling. Its default of four keeps headroom for the coordinator and OS while still exposing useful parallelism. It cannot protect one machine running several Foremen: three coordinators with a ceiling of four could otherwise launch 12 workers. Every Foreman therefore also reserves from one atomic weighted pool under ~/.babysit/resources/.

With parallel_global_units: auto, the pool uses the smaller of the machine's CPUs and one unit per GiB after a 6 GiB OS reserve, with a minimum of one unit. A configured positive value can lower, but not raise, that host-derived budget. New work also queues while available memory is below 20% or one-minute load reaches 80% of the machine's CPUs. Running workers are never preempted.

ProfileUnitsExclusive host resources
plan1—
standard2—
android-simulator4mobile simulator, GPU
ios-simulator4mobile simulator, GPU
local-ml4GPU

Reservations are global across repositories and Foremen sharing the same BABYSIT_HOME. They are keyed by Foreman and Orca Task, making a retry idempotent. The broker enforces the per-Foreman worker ceiling atomically as well as the global weighted budget. Both reserve and status, and the detached watcher, recover leases across all owners without waiting for a dead Foreman:

  • Terminal Dispatches release capacity immediately.
  • Agents proven exited by Orca's fleet view are stopped by exact Dispatch id; capacity is reclaimed only after settlement is confirmed.
  • Reservations with no new Dispatch are reclaimed after ten minutes if the owner's heartbeat is stale or its record is missing. Resume must heartbeat and repeat reserve immediately before launching, saving the returned id.
  • Live workers survive owner interruption and laptop sleep. Unverifiable workers remain held, with RESOURCE_HELD diagnostics; contact loss alone cannot safely authorize another worker in their slot.

Recovery prints RELEASED_LEASE for each reclaimed slot. Reconciliation has a 15-second overall deadline and two-second per-command deadlines; Orca probes never hold the admission lock. A persisted probe cursor rotates past slow workers so they cannot starve later leases on every tick. The OS releases that lock on process exit, including SIGKILL. Replacement leases have new ids, protecting them from late cleanup of the previous attempt. Old mkdir lock directories are ignored by the new broker; do not run old and new broker binaries concurrently during upgrade.

bash
bbs foreman resource status
bbs foreman resource reserve fm-project \
  --ticket bs-child --task orca-task-id --profile ios-simulator
bbs foreman resource release rsc-0123456789abcdef

ADMISSION=queued is backpressure, not a failed Task. Foreman may dispatch other admitted work and retries the queued Task on its next reconcile tick.

Project review and delivery evidence

Foreman prepares one parent plan, design/prototype and proposed ticket map before creating child worktrees or dispatching production work. Review that project checkpoint once; Foreman reviews the child plans against it. Product scope/design changes return to the parent checkpoint. Approval is bound to the artifact revision, so changed documents require a new review.

To delegate the human design reviews too, explicitly pass --auto:

text
Claude Code  /bbs:foreman --auto <large project>
OMP          /foreman --auto <large project>
Codex        $bbs:foreman --auto <large project>

bbs foreman spawn fm-project --auto and direct adoption with --auto persist the choice on that Foreman record across resumes. It still creates and reviews the artifacts; hold/grant bounds, non-delegable decisions, code review, QA, and finish: authorization remain in force. Existing records without auto: true use human project review; child plan autonomy is unchanged.

bbs foreman report <parent-ticket> reads the last durable report.md, even after Orca closes. Each reconciliation writes the observation time, meaningful ticket titles, worker/phase, branch/head, gate evidence, PR/merge state and cleanup/blockers. It is explicitly a saved snapshot; ask the running Foreman to check status for a fresh reconciliation. Execution, delivery and cleanup are separate: PR_READY does not mean merged, and LANDED_LOCAL does not mean pushed. Missing evidence stays UNKNOWN.

Final project integration QA runs after finish handlers and before Foreman reports done: on the actual landed <base> for finish: land, or a retained qa/<parent> branch composed from verified heads for pr/review. It records the tested branch/SHA and acceptance evidence on the parent. Independent code-bearing tickets still require this final check. QA branch preparation and restoration never reset local base; surface compose/revert are only for the earlier scratch/per-ticket lifecycle. Failures return to child repair and verification. The full protocol is in the Foreman project contract.

Which coding agent runs the work

Configure enabled coding agents and the default agent in Orca. The dashboard's Foremen page shows agent ownership guidance (#/settings redirects there); no BBS form or config file controls agent preferences. Foreman uses Orca discovery on the destination host when starting new workers. Existing workers and foremen resume with their recorded agent/model rather than today's default.

bbs agent detect --json identifies the current harness with evidence; bbs agent list --json reports supported CLIs and installed paths. Detection uses BABYSIT_CURRENT_AGENT, the nearest recognized parent process, then native session markers. It does not query Orca, installed binaries, credentials or config; standalone skills remain usable without Orca. bbs agent resolve is removed; configure new worker routes in Orca.

Direct skill invocation is the normal entrypoint:

text
Claude Code  /bbs:foreman <large project>
OMP          /foreman <large project>
Codex        $bbs:foreman <large project>

The skill's first operation is bbs foreman adopt <id>. It detects the agent, reads and renames the active Orca terminal, records the actual agent dialect, and makes dashboard/watchdog wakes addressable. Re-adoption after compaction is idempotent. It refuses to steal an id, terminal, or repo already bound to a different Foreman. bbs foreman spawn is optional convenience and recovery, not the required launch path. Direct invocation inherits the current CLI's permission mode, so configure that session for unattended tool use before leaving a multi-day run.

New Foreman terminals use an explicit --agent when requested, otherwise the current harness. New worker routes are owned by destination-host Orca facts; old BBS worker/foreman agent/provider/model/effort settings no longer affect launches. Legacy YAML bytes remain untouched; each launch command emits one warning when such values are present, and bbs config set rejects new writes. Providers stay in agent-native configuration unless Orca advertises a supported explicit contract.

Adding an agent is a registry entry in internal/agent, which owns the binary name, the flag that suppresses tool approval, how (or whether) a conversation can be given a durable handle, and how that agent namespaces babysit's skills. Two things it does not own:

  • Skills must be reachable by that agent, and a binary on PATH is not that. Each CLI has its own mechanism, and getting it wrong is silent — the worker launches fine and then cannot resolve its prompt:

    agenthow it finds babysit's skillsprompt shape
    claudethe plugin marketplace/bbs:autopilot
    grokgrok plugin install https://github.com/2found/2build/bbs:autopilot
    ompomp config set skills.customDirectories '["$HOME/.claude/plugins/marketplaces/babysit/.claude/skills"]'/autopilot
    codexcodex plugin marketplace add 2found/2build && codex plugin add bbs@babysit$bbs:autopilot
    cursormake the skills available under .cursor/skills or .agents/skills/autopilot

    bbs foreman worker-command preflights the binary and names the per-agent fix; the install itself is on the operator. Pass --skill autopilot rather than writing the prompt by hand — omp reaches its skills through a flat directory list, so they have no bbs: namespace, while Codex uses a $ sigil. A hard-coded /bbs:autopilot resolves incorrectly for both. (omp plugin install <git-url> looks like the fix and is not: it reports success under --dry-run and then fails for real, being an npm-shaped installer rather than a plugin store.) A flat list also bounds what skill:// can address: one skill directory, no ... Anything a skill names outside its own directory — the shared references/ one level up, a sibling skill, a pack-level doc — is silently retargeted inside the skill (skill://qa/../references/worktrees.md becomes skill://qa/references/worktrees.md) and dies as File not found. SKILL.md files therefore say to read those by path, and tests/test_skill_reference_links.sh guards both halves: that the targets resolve, and that every skill which names one says how to read it.

    • A foreman's session is pinned to the agent that minted it. spawn records it and uses it on resume; contradictory explicit --agent requests refuse. Recorded model, effort, and historical provider values remain available for exact-session recovery. New sessions do not read retired BBS preferences.

    What the handle is depends on the agent, and only claude and grok can be told to use one we chose: they take --session-id <uuid>. omp has no such flag, so a foreman on omp gets a private session directory (--session-dir) and resumes with --continue — unambiguous because nothing else writes to that directory. codex and cursor have neither, so either closed foreman starts a fresh conversation and cold-resumes from ticket + Orca state. It deliberately does not use repo-wide resume --last: several foremen may share one repo, and "last" could attach the wrong project's goal. A uuid is never recorded against an agent that cannot be told to use it.

    For multi-day runs, schedule bbs foreman ensure <id> to recreate a missing Orca terminal. The watcher needs no scheduling: adopt and spawn start a detached bbs foreman watch automatically. The unscoped watcher takes one global flock, while a scoped watch <id> takes an id-specific flock, so repeat check-ins are no-ops without blocking other foremen. The watcher exits on its own once no foreman has an open Orca terminal. bbs foreman watch <id> --once remains available to refresh an idle one by hand. Both prompts reload the Foreman skill and carry --foreman-id <id> so compaction or a cold start cannot erase coordinator identity.

grok needs the directory trusted first. grok keeps a per-folder trust record in ~/.grok/trusted_folders.toml, and it is separate from permission_mode — with always-approve set globally, a first run in an unlisted directory still stops on "Do you trust the contents of this directory?", which --always-approve does not answer. An unattended worker parked on that prompt reads as a hung ticket. Both spawn paths preflight it and refuse with the fix named, so the failure is loud instead of silent. Grant it once per repo:

bash
cd <repo> && grok      # answer the trust prompt, then quit

Workers launch with --cwd <repo>, so this is one decision per repo, not per worktree.

Provider, model and effort

Orca owns enabled-agent and default-agent selection for new workers. Foreman applies the phase/tier routing table and destination-host discovery. Explicit per-dispatch --agent, --model, and --effort requests remain available through route and worker-command; managed spawn also accepts explicit fields for a new Foreman. Provider selection belongs in the agent's native configuration unless Orca advertises a supported selector. BBS global and role-specific environment fallbacks are retired; bbs config set rejects retired preferences.

bash
bbs foreman worker-command --agent omp --model '<your-model-id>' --effort high \
  --skill autopilot --prompt 'Build the ticket'
bbs foreman spawn fm-demo --agent codex --model '<your-model-id>'

OMP worker routes return launchMode: native-terminal: render an idle command with bbs foreman worker-command --startup-only --agent omp --model <selection>, create an Orca terminal with that exact command, then call orca orchestration worker-start --terminal <handle> without launch overrides. --startup-only rejects prompts and skill selectors, so supervision delivers the assignment once. This preserves native aliases such as @slow without changing global OMP configuration or falling back to the native default.

Managed Foreman records pin the selected agent, model and effort. Provider values recorded by earlier versions remain readable and are applied only when recovering that pinned session; new launches use native provider configuration. Changed explicit settings are refused for a pinned session.

Foreman routes new supervised workers through bbs foreman route using the destination host's Orca facts and persists the selected route. Receipt verification reads dispatchId and launch.effective.agent, and requires ready/accepted input. For a native OMP terminal, route verify --terminal <handle> returns transport-matched; null effective settings stay null. Verify the worker's actual session settings separately before accepting phase work. After settlement, release the worker and close a retained native terminal only if this Foreman created it. Direct worker-command uses the same route policy. See model routing and the worker launch reference.

Telemetry

Skill runs append JSON Lines to ~/.babysit/analytics/skill-usage.jsonl. Because babysit runs unattended, telemetry is the primary feedback channel — treat it as load-bearing, not decoration. Local-only by default; nothing leaves the machine.

Nothing summarizes that file on demand — the reader is the /bbs:analytics-review skill. Dispatch it by hand (/bbs:analytics-review) when you want a report; to look at the raw rows, read the JSONL directly.

The companion file ~/.babysit/analytics/decisions.jsonl is the audit trail for ticket-level events (size resizes, approval self-resolves) — one line per decision. Grep or jq it directly.

Auto-update

bbs update check compares the local VERSION against main on GitHub, with cache-friendly TTLs (60 min when up-to-date, 12 h when an upgrade is pending). Typical preamble wiring:

bash
UPD="$(bbs update check 2>/dev/null || true)"
case "$UPD" in
  "UPGRADE_AVAILABLE "*) echo "babysit update available — run bbs update";;
  "JUST_UPGRADED "*)     echo "babysit upgraded: $UPD";;
esac

Snooze a pending update: bbs update --snooze 1 (24 h), 2 (48 h), 3 (7 d).

Workflow linting

Every workflow file must declare needs-state: frontmatter so the autopilot orchestrator can route mechanically. bbs autopilot lint-workflow <path> validates this and checks for missing > produces: directives.

bash
# Lint a single workflow
bbs autopilot lint-workflow .claude/skills/autopilot/workflows/builder.md

# Lint all workflows
for wf in .claude/skills/autopilot/workflows/*.md .claude/workflows/*.md; do
  [ -f "$wf" ] && bbs autopilot lint-workflow "$wf"
done

Pre-commit hook

bbs setup installs a pre-commit hook that auto-lints staged workflow files. To install or reinstall:

bash
go run ./cmd/bbs setup

CI

The Lint Workflows GitHub Action runs on pushes and PRs that touch workflow .md files. See .github/workflows/lint-workflows.yml.

Watching a foreman

A foreman drives its batch from its own terminal, so its worst failure is the quiet one: the session finishes a thought, prints nothing more, and sits at an idle prompt while its workers wait for a design gate. Nothing detects that today — the record's heartbeat is written by the foreman itself, so a foreman that stopped working also stopped reporting that it stopped.

bbs foreman watch is the outside observer. It captures the last N lines of the foreman's Orca terminal on an interval; if those bytes are identical for longer than --idle, it types a nudge into the pane — the same "check status" a human would send — and if the nudges stop landing, it says so and gives up rather than poking forever. Independently of the pane, every --status-interval it sends the same skill prompt as an active status check, so a busy foreman still gets asked; a foreman whose record says done leaves the watch set even while its terminal stays open. The status interval defaults to the configured foreman_status_interval (3600 seconds). The idle threshold defaults to that same interval; --status-interval overrides it for that watcher and --idle can independently override the idle threshold. Foreman blocks awaiting worker reports between reminders. Empty check --wait timeouts only renew the wait; they do not trigger project audits. Worker reports trigger focused verification of affected tickets, while status reminders trigger full reconciliation.

bash
bbs foreman watch                       # every foreman with an open workspace
bbs foreman watch fm-acme               # just this one
bbs foreman watch --idle 300 --nudge "check status and report the board"
bbs foreman watch --once                # one pass, for cron
flagdefaultwhat it does
--interval <sec>60how often to capture the pane
--idle <sec>effective --status-interval (3600)unchanged for this long → nudge
--status-interval <sec>foreman_status_interval config (3600)periodic status prompt, even while the pane moves
--lines <n>40how much of the pane forms the fingerprint
--nudge <text>check statuswhat gets typed in
--max-nudges <n>3budget before it reports STALLED and stops
--onceoffsingle pass then exit; prints a line per foreman

It is a foreground loop, not a daemon — but you rarely start it yourself: bbs foreman adopt and bbs foreman spawn launch an unscoped detached copy on every check-in. Its global flock keeps exactly one unscoped watcher running; scoped watchers use separate per-foreman flocks and cannot block other foremen. It writes only its own clock and watch.log under ~/.babysit/watch/, and exits when no foreman has an open Orca terminal. Output is events only — a foreman that is working produces no output at all.

text
NUDGED fm-acme after 12m (1/3) — sent "check status"
STATUS fm-acme after 60m — sent "check status"
STALLED fm-acme — 3 nudges, no change in 41m; open "bbs foreman"
GONE fm-acme — terminal "bbs foreman" is closed

Two behaviours worth knowing. It selects foremen by open Orca terminal, not by liveness — a foreman wedged long enough to need a nudge is exactly the one whose heartbeat has gone stale, so selecting on Live() would drop every foreman this exists to catch. And the nudge's own echo in the pane does not refund the budget: real progress changes the pane on more than one tick, which is what keeps --max-nudges binding on a dead session. The status clock is separate from the idle clock on purpose: a status prompt neither spends a nudge nor resets the idle window, so it can never let an unresponsive terminal slip past the stall bound.

Health checks

There is no health-check command. bbs setup reports what it linked and warns when ~/.local/bin is missing from your PATH; beyond that, the preamble is the live check — it emits BBS_DEGRADED on stderr at the top of every skill run when no working bbs is reachable, which is the failure that actually matters.

To verify an install by hand:

bash
bbs ticket --help     # exit 0 = the binary is present and serving subcommands
bbs config list       # reads ~/.babysit/config.yaml