Pier is a Harbor-compatible framework for evaluating coding agents in sandboxed environments. It reads Harbor's task format and runs trials against it.
pier run -p path/to/task --agent claude-code --env modalPier is a fork. We wanted a smaller, more opinionated base to build on. On top of Harbor, Pier adds:
- Installed agents in air-gapped tasks (
allow_internet = false). When the agent runs inside the sandbox (Claude Code, Codex, etc.), both the install step and the inference call need the network. Pier lets agents declare their install scripts and a network allowlist, whichdockerandmodalenvironments honor when setting up the sandbox. - Augmented ATIF v1.7. Strict one step per API turn, strict reasoning vs agent message separation, no fabricated assistant text,
peak_context_tokens,summarization_count,llm_call_count, real upstream timestamps. - A chat-style trajectory viewer (
pier view). pier critique runfor inspecting completed trials with a fresh agent in a fresh sandbox.
- Task format: Harbor-compatible.
- Environments:
docker,modal. Per-agent install specs and network allowlists are honored on both, so installed agents work underallow_internet = false. - Agents:
nop,oracle,antigravity-sdk,claude-code,codex,cursor-cli,gemini-cli,opencode,mini-swe-agent,pi. All emit augmented ATIF v1.7. - Datasets: local Harbor-format task directories via
-p/--path. - CLI:
pier run,pier job,pier view,pier critique run,pier check/pier analyze(vendored from Harbor)
Pier does not currently resolve or download Harbor registry datasets directly.
uv tool install datacurve-pier
# or
pip install datacurve-pierexport ANTHROPIC_API_KEY=...
pier run -p path/to/task --agent claude-code --env modal --env-file .envRun a local dataset, optionally a deterministic random subset:
pier run -p path/to/dataset --agent claude-code --env modal
pier run -p path/to/dataset --n-tasks 10 --sample-seed 0To use a Harbor registry dataset, download it with Harbor first, then point Pier at it:
uv run --directory ~/code/harbor harbor download swebenchpro -o ~/code/pier/datasets
uv run pier run -p datasets/swebenchpro --n-tasks 10 --sample-seed 0Trials land under jobs/<timestamp_or_name>/<trial_id>/. See pier run --help, pier job --help, pier critique --help, and pier view --help for everything else.
If your inference gateway is served with an internal CA, point PIER_EXTRA_CA_CERTS at a PEM bundle on the host:
PIER_EXTRA_CA_CERTS=/path/to/internal-roots.pem pier run -p path/to/task --agent codexPier appends a root install step to every installed agent that registers those certificates with the sandbox's OS trust store, and exports the CA variables each runtime reads: SSL_CERT_FILE, REQUESTS_CA_BUNDLE, CURL_CA_BUNDLE and GIT_SSL_CAINFO get a bundle of the distro roots plus yours (these replace the default store), while NODE_EXTRA_CA_CERTS and CODEX_CA_CERTIFICATE get your certificates alone (these add to it). Codex needs CODEX_CA_CERTIFICATE specifically -- the npm package is only a launcher shim around a static Rust binary, so NODE_EXTRA_CA_CERTS does nothing for it.
Because the replace-style variables substitute a runtime's trust store rather than adding to it, the merged bundle unions the public sources it can find -- every readable distro bundle, plus certifi for each of a bounded list of interpreters (the system ones, Pier's agent venv, and uv tool environments under /root and /home/*) -- so a Python agent normally trusting a newer certifi does not lose roots. That list is a guess, not an enumeration: a Python runtime installed somewhere unprobed, holding certifi roots none of those interpreters have, could still see its trust narrowed by REQUESTS_CA_BUNDLE. This does not affect the Node, Bun or Rust harnesses. Agent installs run before the CA step and are given no CA variables at all, since they fetch from public hosts and a variable pointing at a not-yet-written file is a hard failure. Bare debian:12 ships no trust store; if no public roots are found anywhere the step warns in the build log. The certificates are baked into the image at build time and counted in the install fingerprint, so rotating the CA invalidates cached sandbox images. Works on both docker and modal, and under allow_internet = false -- the egress proxy tunnels HTTPS with CONNECT and never re-signs the origin certificate, so the agent validates the gateway's real cert in every network mode.
The variables are injected as defaults; an explicit agent.env entry for the same key still wins.
Both halves are needed because the runtime that makes the request differs per agent, and even per agent version:
| Runtime | Reads the OS trust store? | Needs a variable |
|---|---|---|
| Rust (codex) | yes, via rustls_native_certs |
CODEX_CA_CERTIFICATE also honoured |
| Bun native binary (opencode, claude-code) | yes | NODE_EXTRA_CA_CERTS / SSL_CERT_FILE also honoured |
| Node JS via npm (pi, gemini-cli) | no -- Node ships its own root list | NODE_EXTRA_CA_CERTS required |
| Python (mini-swe-agent, antigravity-sdk) | via SSL_CERT_FILE; requests/httpx default to certifi |
SSL_CERT_FILE, REQUESTS_CA_BUNDLE |
Node is the reason the OS trust store alone is not sufficient: with a private root correctly installed via update-ca-certificates, curl succeeds while Node 22 fails with UNABLE_TO_VERIFY_LEAF_SIGNATURE (tls.rootCertificates is a separate, bundled list). Conversely a Bun-packaged agent works from the trust store with no variables at all -- which is why Pier does both rather than choosing.
Verified end to end against a server signed by a private root: codex 0.151.0 (CODEX_CA_CERTIFICATE), opencode 1.18.25 (trust store alone, and each variable independently), and Node 22 (NODE_EXTRA_CA_CERTS). The tests additionally assert that every registered agent receives the install step and the variables. When adopting a new harness, smoke-test that both private and public HTTPS succeed inside the sandbox rather than assuming its packaging matches a sibling's.
Use agent.model_name for trial metadata, agent.env for runtime env vars, and agent-specific kwargs for tool config. Pier's network allowlist also reads URLs out of those configs (Codex config_toml, OpenCode opencode_config, mini-swe config_yaml), so any base URL you set is allowlisted without code changes.
A few things we've learned plumbing this through Respan and OpenRouter:
Claude Code routes through the Anthropic face from Respan. Plan mode is disabled by default (--disallowedTools EnterPlanMode).
- name: claude-code
model_name: claude-opus-4-7
env:
ANTHROPIC_AUTH_TOKEN: ${RESPAN_API_KEY}
ANTHROPIC_BASE_URL: https://endpoint.respan.ai/api/anthropic
ANTHROPIC_CUSTOM_HEADERS: "X-Respan-Route-Provider: vertex_ai"
kwargs:
reasoning_effort: maxCodex needs a [model_providers.<name>] block with wire_api = "responses" (not WebSockets, which Codex defaults to and Respan doesn't speak).
- name: codex
model_name: openai/gpt-5.5
env: { RESPAN_API_KEY: ${RESPAN_API_KEY} }
kwargs:
config_toml: |
model_provider = "respan"
[model_providers.respan]
name = "Respan Gateway"
base_url = "https://endpoint.respan.ai/api/"
wire_api = "responses"
env_key = "RESPAN_API_KEY"
reasoning_effort: xhighGemini CLI:
- name: gemini-cli
model_name: gemini/gemini-3.1-pro-preview
env:
GEMINI_API_KEY: ${RESPAN_API_KEY}
GOOGLE_GENERATIVE_AI_API_KEY: ${RESPAN_API_KEY}
GEMINI_API_BASE: https://endpoint.respan.ai/api/google/vertexai/v1beta
GOOGLE_GEMINI_BASE_URL: https://endpoint.respan.ai/api/google/vertexai/Antigravity SDK runs Google's Python SDK with its platform-specific local
harness in an isolated Python 3.12 environment with hash-verified, fully locked
dependencies. It supports Pier skills, stdio and streamable-HTTP MCP servers,
and live ATIF checkpoints. reasoning_effort accepts minimal, low, medium,
or high; None uses medium. SSE MCP servers are not supported by
google-antigravity 0.1.9.
- name: antigravity-sdk
model_name: google/gemini-3.6-flash
env:
GEMINI_API_KEY: ${GEMINI_API_KEY}
kwargs:
reasoning_effort: highCursor CLI uses the installed cursor-agent binary, so it fits the same
inside-the-sandbox path as Claude Code, Codex, Gemini CLI, and OpenCode. Use
cursor/composer-2.5 for Composer 2.5 trial metadata and pass CURSOR_API_KEY
through your env file.
- name: cursor-cli
model_name: cursor/composer-2.5
env:
CURSOR_API_KEY: ${CURSOR_API_KEY}OpenCode uses opencode_config to add unknown providers or override known ones. To redirect Google to Respan, override just options.baseURL; to add a fully custom provider, use opencode_config.provider.<name> with the npm package, options, and models.
Pi resolves --model against its own built-in catalog, so a gateway or proxy needs a provider entry in models.json. Setting the provider's base-URL env var (OPENAI_BASE_URL, ANTHROPIC_BASE_URL) generates one automatically, which routes pi's built-in models through the endpoint while keeping their shipped cost and capability metadata. Use pi_config to declare a slug that is not in the catalog; supply its cost (per million tokens) so pi can price it, otherwise Pier falls back to the LiteLLM price table and leaves cost_usd unset for a private slug. pi_config is deep-merged over the generated config, and any baseUrl in it is added to the network allowlist.
Pi has no built-in MCP support (configured MCP servers are ignored with a warning) and no resume support. Skills are copied to ~/.agents/skills. PI_OFFLINE=1 and PI_SKIP_VERSION_CHECK=1 are set by default so pi's startup update check does not hit the egress proxy on no-network tasks; override them via env if needed.
- name: pi
model_name: openai/my-proxy-slug
env:
OPENAI_API_KEY: ${OPENAI_API_KEY}
OPENAI_BASE_URL: ${OPENAI_BASE_URL}
kwargs:
thinking: medium
pi_config:
providers:
openai:
api: openai-completions
models:
- id: my-proxy-slug
reasoning: true
contextWindow: 400000
cost: { input: 1.25, output: 10.0, cacheRead: 0.125 }mini-swe-agent picks a native adapter from the model-name prefix: openai/... → litellm_response (OpenAI Responses end-to-end), openrouter/... → openrouter (BYOK costs from cost_details.upstream_inference_cost), everything else → LiteLLM auto.
For Gemini 3 via mini-swe-agent/LiteLLM, omitting reasoning_effort uses the Gemini API default high/dynamic thinking level, but it does not request readable thought summaries. Set kwargs.reasoning_effort: high explicitly when you want LiteLLM to send includeThoughts and preserve returned summaries as reasoning content.
- name: mini-swe-agent
model_name: openrouter/qwen/qwen3.6-plus
env: { OPENROUTER_API_KEY: ${OPENROUTER_API_KEY} }
kwargs:
set_cache_control: default_end