Astvyr
BENCHMARKS AND SLO

Targets published so they can be checked.

These are engineering targets measured at reference conditions, not marketing numbers. Where a target is not met yet, this page says so.

PRODUCT TARGETS

What the product is held to.

Measured on reference hardware, reference repositories and named regions. A target without a measurement condition is not a target.

< 20 minTime to first verified changeOn a supported starter repository
> 70%Accepted mission rateWithout a full manual rewrite
< 0.5%False green rateOn monitored production missions
50-80%Lead time, intent to verified releaseImprovement across selected task classes
99.99%Platform availabilityEnterprise control plane, with graceful local mode
> 70%Gross marginBlended, after scale; compute billed transparently
> 125%Net revenue retentionEnterprise; team expansion driven by outcome
ZeroPolicy bypass incidentsAny occurrence leads to a release stop
PERFORMANCE

Latency and responsiveness targets.

Every row names the measurement condition, because a p95 without a context is a slogan.

MetricTargetMeasurement condition
Desktop startup, warm< 2 s p50Reference hardware and workspace
Desktop startup, cold< 5 s p95Reference hardware and workspace
Editor input and navigation< 50 ms p95Local telemetry, without content
Inline completion, first suggestion< 500 ms p95End to end, by region and model
Mission UI events< 2 s p95Server event to active client
Repository search< 1 s p95Reference monorepo, after indexing
Initial index, 1M LOC< 5 min partial, < 20 min completep95, reference languages and hardware
Warm sandbox ready< 30 s p95Standard image, regional capacity
Production Twin, basic environment< 10 min p95Documented reference stack
Policy decision< 100 ms p95Service-side, cached input, regional
Release pause signal applied< 30 s p95Detection to controller command
AVAILABILITY

Availability and integrity, per system.

Local editing availability is deliberately independent of the cloud: a bad day for a provider is not a bad day for your editor.

SystemTargetNote
Desktop local editing99.99%Independent of the cloud
Regional API and control plane99.95%Monthly; higher tiers can contract an enhanced target
Mission orchestration99.95%Durable availability without losing state
Authentication99.99%With emergency cached access for local work
Release control99.99%Active rollout control on reserved capacity
Audit durabilityNo acknowledged event lossMulti-copy with reconciliation and alerting
Evidence integrityNo undetected mutationHashes and signatures with periodic verification
Object durability≥ 11 ninesCloud-provider class where available
RECOVERY

Recovery objectives, by failure scope.

Drills are scheduled against these numbers: regional failover at least twice a year on a mature cell, database restore monthly and quarterly.

ScopeRecovery pointRecovery time
Single service or zoneNear zero, event replayMinutes
Regional cell≤ 5 min for transactional metadata≤ 60 min for priority services
Global directoryNear zero≤ 30 min
Object evidence storeVersioned and replicatedHours for bulk access; critical manifests first
Customer runner lossNo control-plane state lostRe-enrolment at customer capacity
Signing key compromiseNot applicableEmergency chain within hours
VERIFIED DELIVERY BENCHMARK

A benchmark that includes our failures.

The public benchmark is reproducible by outsiders, uses hidden tests and publishes false-green rates and cost — including the runs where Astvyr lost.

Mission success rate

Accepted outcomes divided by eligible Missions, tracked by cohort and task class.

False green rate

Declared safe or successful, but failed acceptance or caused an escape. The key trust KPI.

Human rework ratio

How much work a person has to redo, tracked without weakening outcomes.

Retrieval critical recall

Must stay above the threshold for high-risk tasks.

Tool error and denial rate

Denial can be a healthy control, so it is read together with errors.

Cost per verified outcome

Total model and compute cost of an accepted Mission, which must fall with scale.

Time to verified outcome

From an approved contract to evidence-ready.

NOT OPTIMISED IN ISOLATION
  • Lines of generated code without an accepted outcome.
  • Number of agent tool calls without efficiency and safety.
  • Mission execution time without accounting for false green and rework.
  • Daily active users bought with free inference without unit economics.
  • Autonomy as a percentage of actions without human approval.
QUALITY

Ten layers of checking before a release gate.

  • Component: unit, property, schema and deterministic state tests.
  • Service: integration with databases, queues, identity and policy, plus failure injection.
  • Contract: public APIs, events, provider and tool adapters, desktop and server compatibility.
  • System: Mission, sandbox, evidence, release and incident journeys.
  • Desktop matrix: Windows, macOS and Linux, upgrades, extensions, remote workspaces, accessibility.
  • Security: SAST, SCA, IaC and container scanning, threat tests, penetration tests, red team, sandbox escape programme.
  • AI evaluation: task, retrieval, tool, adversarial, false-green and outcome regression.
  • Performance: latency, indexing, concurrency, scheduler, storage and regional failover.
  • Resilience: chaos, dependency outage, provider degradation, queue backlog and restore drills.
  • Compliance evidence: control operation, audit completeness, retention and access reviews.
RELEASE GATES
  • No unresolved critical security finding, unless an executive and security approval records a compensating control.
  • No regression beyond the error budget in tests, AI evaluations, performance or accessibility.
  • The migration is tested from every supported version, with a documented rollback and downgrade.
  • Telemetry, dashboards, alerts, runbooks, capacity and support training are ready before the feature flag widens.
  • Privacy, data flow and the threat model are updated for every new data source, model, tool or integration.
  • Staged rollout: internal, design partners, percentage canary, region waves, then general availability.
NORTH STAR

One metric the whole product answers to.

Everything below decomposes into this number. A metric that cannot be broken into owned branches is a slogan.

NORTH STAR METRIC

Verified Production Outcomes per active team per month

Missions that passed the approved criteria, independent evidence and the observation window in the real intended environment: changes that passed the mandatory checks, were delivered by policy and were not rolled back because of a product error inside the observation window. Work without production uses a Verified Accepted Outcome, counted separately.

METRIC TREE

Ten branches, from acquisition to ecosystem.

The north star decomposed into the branches that are tracked, and what each branch is read through.

Acquisition

Qualified visitors, installs, source quality, cost, organic share.

Activation

Repository connected, first contract approved, first evidence-ready Mission, time-to-value.

Engagement

Weekly active developers and teams, Missions per user, supported workflows, IDE days active.

Quality and trust

Mission success, false-green, rework, rollback, security denials, satisfaction after failure.

Production value

Lead time, deployment frequency, change failure, MTTR, incident recurrence.

Retention

D1, W4 and M3 developer retention, team logo and seat retention, feature cohort retention.

Revenue

Free-to-paid, ARPA, usage, expansion, NRR, churn, collections.

Economics

Model and compute cost, contribution margin, support cost, CAC payback, burn multiple.

Reliability

SLO, job success, queue delay, restore and failover, support response.

Ecosystem

Active publishers, trusted packages, installs, GMV, package incidents.

EXPERIMENTATION

Five rules that govern any test we run.

Written before an experiment starts, so a result cannot be declared after the fact.

  • An experiment has a hypothesis, a target cohort, a primary metric, guardrails, a duration and a decision owner.
  • Activation or usage is never optimised at the cost of hidden data transfer, unsafe autonomy or surprise billing.
  • AI and model tests are segmented by task and risk; an average number must not hide a critical failure.
  • Enterprise contractual and control changes are not put through random A/B tests without consent and governance.
  • Failed experiments are documented; repeating a proposal requires new data.
EXECUTIVE DASHBOARD

Five questions the leadership review answers.

Each question is answered by indicators, not by narrative.

  • Do we create real value? Verified outcomes, time saved, accepted without rework, production improvements.
  • Can the product be trusted? False-green, escapes, unsafe action attempts, incidents and recovery.
  • Is the habit growing? Cohort retention, weekly IDE use, team invites and workflow breadth.
  • Is the economics healthy? Contribution margin per Mission, model and compute cost, CAC payback, NRR.
  • Is scale durable? SLO, queue and capacity, support load, regional health and security findings.
SCALE

Designed capacity, not a demand forecast.

These are architectural targets that shape early decisions. They are not a claim about demand.

10M+REGISTERED DEVELOPERS
3M+MONTHLY ACTIVE DEVELOPERS
250K+ACTIVE ORGANISATIONS
500K+CONCURRENT IDE SESSIONS
5M+MISSIONS PER DAY
200K+CONCURRENT SANDBOXES
petabyte classARTIFACTS PER DAY
Americas, Europe, APAC + sovereignREGIONS
MEASUREMENT

If it cannot be reproduced, it is not a benchmark.

Hidden tests, published cost, false-green accounting and honest failures. A number nobody else can reproduce is marketing.