Private, maintainer-authored evals

Evaluate coding agents where software actually runs.

Olympus carries frontier agents beyond the pull request and into deployment, telemetry, incidents, recovery, regression safety, and real operational consequences.

Private codebase Resettable infrastructure Original maintainer grading
olympus / run-04 LIVE
COMPOUND DEPLOYMENT TASKMODEL ROLLOUT 4/4
feature_implemented+18m
tests_passed+31m
deployed_to_cluster+39m
sev_2: queue saturation+47m
mitigation incomplete+63m
CONTINUOUS GRADER0.43

Meaningful failure: recovery restored throughput but violated data integrity.

The environment

Static tasks stop where production judgment begins.

Olympus is a production software and infrastructure platform spanning distributed services, workflow execution, Postgres, Temporal, Kubernetes, and multi-cloud operations.

01

Implement

The agent works inside an unfamiliar, private codebase with real architectural constraints.

02

Deploy

The change moves into an isolated runtime with services, workflows, databases, and clusters.

03

Operate

Telemetry, load, degraded dependencies, and reproducible incidents force live decisions.

04

Grade

Results reflect correctness, mitigation, recovery, regressions, cost, and engineering judgment.

The difference

The engineer who built the system authors the evaluation.

That creates tasks around the failure modes, tradeoffs, and operational consequences that are usually invisible to external task writers.

Repository-only evaluation

Dead code and proxy judgment

  • Ends at tests or patch output
  • Task authors did not build the system
  • Operational consequences are simulated in prose
  • Binary graders hide partial and compound failures
Olympus live evaluation

Running systems and maintainer judgment

  • Continues through deployment and runtime behavior
  • Authored by the platform’s original engineer
  • Reproduces Sev incidents with bounded cloud impact
  • Captures multiple meaningful failures per rollout
Paid pilot

Prove discrimination before scaling.

A paid qualification validates task difficulty, environment compatibility, and grading quality before either side commits to the complete pilot.

Review the qualification
Q0

Qualification

Validate environment integration and meaningful model discrimination.

GATE
E1

Production feature and deployment

Implementation quality under live system constraints.

4+ ROLLOUTS
E2

Reproducible Sev incident response

Diagnosis, mitigation, recovery, and regression safety.

4+ ROLLOUTS
E3

Compound implementation to recovery

Long-horizon performance across code, deployment, and operations.

4+ ROLLOUTS
Grading method

Measure how the run fails, not only whether it passes.

Continuous scores separate superficial completion from engineering-quality outcomes. A single rollout can expose multiple documented failure modes, including changes a senior engineer would reject despite passing tests.

ROLLOUT SCORE0.00 TO 1.00
0.000.500.901.00
Service restored, data integrity violated0.43
Feature works, unsafe rollback path0.58
Correct patch, excessive cloud cost0.71
Controlled execution

Live enough to matter. Bounded enough to repeat.

Each evaluation runs in an isolated environment designed for reset, reproducibility, controlled credentials, and bounded infrastructure consequences.

01

Private by design

The Olympus codebase and evaluation materials remain confidential and access-controlled.

02

Resettable state

Services, databases, workflows, and incident conditions return to a known baseline between runs.

03

Bounded consequences

Cloud access, credentials, spending, and operational impact are constrained per evaluation.

Olympus by Zeusfyi

Put a frontier coding agent inside a system that can push back.

Review the private environment, qualification criteria, and proposed three-eval pilot with the engineer who built Olympus.