Targets published so they can be checked.
These are engineering targets measured at reference conditions, not marketing numbers. Where a target is not met yet, this page says so.
What the product is held to.
Measured on reference hardware, reference repositories and named regions. A target without a measurement condition is not a target.
Latency and responsiveness targets.
Every row names the measurement condition, because a p95 without a context is a slogan.
| Metric | Target | Measurement condition |
|---|---|---|
| Desktop startup, warm | < 2 s p50 | Reference hardware and workspace |
| Desktop startup, cold | < 5 s p95 | Reference hardware and workspace |
| Editor input and navigation | < 50 ms p95 | Local telemetry, without content |
| Inline completion, first suggestion | < 500 ms p95 | End to end, by region and model |
| Mission UI events | < 2 s p95 | Server event to active client |
| Repository search | < 1 s p95 | Reference monorepo, after indexing |
| Initial index, 1M LOC | < 5 min partial, < 20 min complete | p95, reference languages and hardware |
| Warm sandbox ready | < 30 s p95 | Standard image, regional capacity |
| Production Twin, basic environment | < 10 min p95 | Documented reference stack |
| Policy decision | < 100 ms p95 | Service-side, cached input, regional |
| Release pause signal applied | < 30 s p95 | Detection to controller command |
Availability and integrity, per system.
Local editing availability is deliberately independent of the cloud: a bad day for a provider is not a bad day for your editor.
| System | Target | Note |
|---|---|---|
| Desktop local editing | 99.99% | Independent of the cloud |
| Regional API and control plane | 99.95% | Monthly; higher tiers can contract an enhanced target |
| Mission orchestration | 99.95% | Durable availability without losing state |
| Authentication | 99.99% | With emergency cached access for local work |
| Release control | 99.99% | Active rollout control on reserved capacity |
| Audit durability | No acknowledged event loss | Multi-copy with reconciliation and alerting |
| Evidence integrity | No undetected mutation | Hashes and signatures with periodic verification |
| Object durability | ≥ 11 nines | Cloud-provider class where available |
Recovery objectives, by failure scope.
Drills are scheduled against these numbers: regional failover at least twice a year on a mature cell, database restore monthly and quarterly.
| Scope | Recovery point | Recovery time |
|---|---|---|
| Single service or zone | Near zero, event replay | Minutes |
| Regional cell | ≤ 5 min for transactional metadata | ≤ 60 min for priority services |
| Global directory | Near zero | ≤ 30 min |
| Object evidence store | Versioned and replicated | Hours for bulk access; critical manifests first |
| Customer runner loss | No control-plane state lost | Re-enrolment at customer capacity |
| Signing key compromise | Not applicable | Emergency chain within hours |
A benchmark that includes our failures.
The public benchmark is reproducible by outsiders, uses hidden tests and publishes false-green rates and cost — including the runs where Astvyr lost.
Mission success rate
Accepted outcomes divided by eligible Missions, tracked by cohort and task class.
False green rate
Declared safe or successful, but failed acceptance or caused an escape. The key trust KPI.
Human rework ratio
How much work a person has to redo, tracked without weakening outcomes.
Retrieval critical recall
Must stay above the threshold for high-risk tasks.
Tool error and denial rate
Denial can be a healthy control, so it is read together with errors.
Cost per verified outcome
Total model and compute cost of an accepted Mission, which must fall with scale.
Time to verified outcome
From an approved contract to evidence-ready.
- Lines of generated code without an accepted outcome.
- Number of agent tool calls without efficiency and safety.
- Mission execution time without accounting for false green and rework.
- Daily active users bought with free inference without unit economics.
- Autonomy as a percentage of actions without human approval.
Ten layers of checking before a release gate.
- Component: unit, property, schema and deterministic state tests.
- Service: integration with databases, queues, identity and policy, plus failure injection.
- Contract: public APIs, events, provider and tool adapters, desktop and server compatibility.
- System: Mission, sandbox, evidence, release and incident journeys.
- Desktop matrix: Windows, macOS and Linux, upgrades, extensions, remote workspaces, accessibility.
- Security: SAST, SCA, IaC and container scanning, threat tests, penetration tests, red team, sandbox escape programme.
- AI evaluation: task, retrieval, tool, adversarial, false-green and outcome regression.
- Performance: latency, indexing, concurrency, scheduler, storage and regional failover.
- Resilience: chaos, dependency outage, provider degradation, queue backlog and restore drills.
- Compliance evidence: control operation, audit completeness, retention and access reviews.
- No unresolved critical security finding, unless an executive and security approval records a compensating control.
- No regression beyond the error budget in tests, AI evaluations, performance or accessibility.
- The migration is tested from every supported version, with a documented rollback and downgrade.
- Telemetry, dashboards, alerts, runbooks, capacity and support training are ready before the feature flag widens.
- Privacy, data flow and the threat model are updated for every new data source, model, tool or integration.
- Staged rollout: internal, design partners, percentage canary, region waves, then general availability.
One metric the whole product answers to.
Everything below decomposes into this number. A metric that cannot be broken into owned branches is a slogan.
Verified Production Outcomes per active team per month
Missions that passed the approved criteria, independent evidence and the observation window in the real intended environment: changes that passed the mandatory checks, were delivered by policy and were not rolled back because of a product error inside the observation window. Work without production uses a Verified Accepted Outcome, counted separately.
Ten branches, from acquisition to ecosystem.
The north star decomposed into the branches that are tracked, and what each branch is read through.
Acquisition
Qualified visitors, installs, source quality, cost, organic share.
Activation
Repository connected, first contract approved, first evidence-ready Mission, time-to-value.
Engagement
Weekly active developers and teams, Missions per user, supported workflows, IDE days active.
Quality and trust
Mission success, false-green, rework, rollback, security denials, satisfaction after failure.
Production value
Lead time, deployment frequency, change failure, MTTR, incident recurrence.
Retention
D1, W4 and M3 developer retention, team logo and seat retention, feature cohort retention.
Revenue
Free-to-paid, ARPA, usage, expansion, NRR, churn, collections.
Economics
Model and compute cost, contribution margin, support cost, CAC payback, burn multiple.
Reliability
SLO, job success, queue delay, restore and failover, support response.
Ecosystem
Active publishers, trusted packages, installs, GMV, package incidents.
Five rules that govern any test we run.
Written before an experiment starts, so a result cannot be declared after the fact.
- An experiment has a hypothesis, a target cohort, a primary metric, guardrails, a duration and a decision owner.
- Activation or usage is never optimised at the cost of hidden data transfer, unsafe autonomy or surprise billing.
- AI and model tests are segmented by task and risk; an average number must not hide a critical failure.
- Enterprise contractual and control changes are not put through random A/B tests without consent and governance.
- Failed experiments are documented; repeating a proposal requires new data.
Five questions the leadership review answers.
Each question is answered by indicators, not by narrative.
- Do we create real value? Verified outcomes, time saved, accepted without rework, production improvements.
- Can the product be trusted? False-green, escapes, unsafe action attempts, incidents and recovery.
- Is the habit growing? Cohort retention, weekly IDE use, team invites and workflow breadth.
- Is the economics healthy? Contribution margin per Mission, model and compute cost, CAC payback, NRR.
- Is scale durable? SLO, queue and capacity, support load, regional health and security findings.
Designed capacity, not a demand forecast.
These are architectural targets that shape early decisions. They are not a claim about demand.
If it cannot be reproduced, it is not a benchmark.
Hidden tests, published cost, false-green accounting and honest failures. A number nobody else can reproduce is marketing.