Test Temp Directories
Use the shared helpers intest/helpers/temp-dir.ts for test-owned temporary
directories so ownership is explicit and cleanup stays in the test lifecycle:
useAutoCleanupTempDirTracker(afterEach) intentionally exposes no manual
cleanup method - Vitest owns cleanup after each test. Older lower-level
helpers (makeTempDir, cleanupTempDirs, createTempDirTracker) still exist
for tests that have not migrated; avoid new usage of them and avoid new bare
fs.mkdtemp* calls unless a test is explicitly verifying raw temp-dir
behavior. When a bare temp dir is genuinely needed, add an auditable allow
comment with a reason:
node scripts/report-test-temp-creations.mjs reports new bare temp-dir
creation and new manual shared-helper usage in added diff lines, without
blocking existing cleanup styles. It follows the same test-path classification
as scripts/changed-lanes.mjs and skips the shared helper implementation
itself. check:changed runs this report for changed test paths as a
warning-only CI signal (GitHub warning annotations, not failures).
Agent reliability evals (skills)
We already have a few CI-safe tests that behave like “agent reliability evals”:- Agent admission, run-ID responses, and abort requests through the real Gateway with a mock OpenAI provider (
src/gateway/gateway.test.ts). - End-to-end wizard flows that validate session wiring and config effects (
src/gateway/gateway.test.ts).
- Decisioning: when skills are listed in the prompt, does the agent pick the right skill (or avoid irrelevant ones)?
- Compliance: does the agent read
SKILL.mdbefore use and follow required steps/args? - Workflow contracts: multi-turn scenarios that assert tool order, session history carryover, and sandbox boundaries.
- A scenario runner using mock providers to assert tool calls + order, skill file reads, and session wiring.
- A small suite of skill-focused scenarios (use vs avoid, gating, prompt injection).
- Optional live evals (opt-in, env-gated) only after the CI-safe suite is in place.
Cost budget
Every test file runs on every pull request that touches its area, so its cost is paid thousands of times. Budgets, measured withpnpm test <file> --maxWorkers=1
on one worker:
- Target under 5 s of test time per file. Above 30 s, the PR body explains which contract needs that time and why no cheaper layer proves it.
- A file that needs more than the planner’s per-job budget (about 120 s) cannot be
packed with other files and sets the wall time of its whole job. Split it by
owner boundary, or move the long end-to-end composition to the release-only tier
(
RELEASE_ONLY_*sets inscripts/lib/ci-node-test-plan.mts) and keep a fast contract check in per-PR CI. - Use an injected clock at the owner instead of real timers, sleeps, or polling;
claim ports through
src/test-utils/port-claims.ts; give each file its own state directory; reuse suite-level Gateway and process fixtures instead of booting per test; import the narrow test API of a plugin or module rather than its full barrel. Do not add a serial Vitest config or a worker pin: fix the shared state that would need one. - State the measured cost in the PR for every new or materially changed test file, and the CI seconds once the run exists.
Flake triage
A failure without a related change is a defect. Never re-run, re-push, or refresh to get green.- Reproduce in the failing shard’s file order first (the plan’s file list is in the job log), then alone. Order-only failures are shared-state leaks from an earlier file.
- Classify: fixture (temp state, ports, cwd, env, module singletons), ordering (assertion before the owned completion signal), or product (a real race).
- Fix at the owner. Product races get a regression that fails on the original defect; fixture leaks are fixed in the fixture owner, not in the failing test.
- Proof bar: 20 clean standalone runs of the file, 3 clean runs of the original shard, and the owner’s sibling tests. Record the root cause in the PR.
- If the owner is another lane or PR, cite that fix; it is the only reason to proceed on a red job.
Adding regressions (guidance)
For inventory-growth and capacity regressions, pass bounded synthetic inventories and fixed timing data through the real planner. Put the fixture at the capacity boundary and preserve coverage, ownership, and execution-budget assertions. Keep real-checkout inventory coverage in its integration tests instead of rebuilding the growing repository plan for every synthetic variation. When you fix a provider/model issue discovered in live:- Add a CI-safe regression if possible (mock/stub provider, or capture the exact request-shape transformation)
- If it’s inherently live-only (rate limits, auth policies), keep the live test narrow and opt-in via env vars
- Prefer targeting the smallest layer that catches the bug:
- provider request conversion/replay bug -> direct models test
- gateway session/history/tool pipeline bug -> gateway live smoke or CI-safe gateway mock test
- SecretRef traversal guardrail:
src/secrets/exec-secret-ref-id-parity.test.tsderives one sampled target per SecretRef class from registry metadata (listSecretTargetRegistryEntries()), then asserts traversal-segment exec ids are rejected.- If you add a new
includeInPlanSecretRef target family insrc/secrets/target-registry-data.ts, updateclassifyTargetClassin that test. The test intentionally fails on unclassified target ids so new classes cannot be skipped silently.