Skip to content

Epic: Harbor-based coding evals to BlindBench review and training data #9

Description

@nmogil

Outcome

Mogil Bench becomes the Harbor-based execution and evidence layer for coding-agent evaluations, with local Docker and Daytona environments, full Pi trajectories, BlindBench human review, and an approval-gated path to open-weights fine-tuning data.

Architecture direction

  • Harbor owns coding task/trial/environment/verifier execution.
  • Mogil Bench owns benchmark packs, model/harness matrices, stable logical identity, evidence policy, and BlindBench workflow.
  • BlindBench owns blinded human review and approved-data reuse.
  • Daytona is an initial remote Harbor environment; local Docker is the first implementation target.
  • Fireworks-compatible SFT/DPO export reuses BlindBench #287; training jobs remain separately approved.

Workstreams

Product gate

The paid-comparison gate was met with fictional/public tasks: 18/18 provider-parity attempts were quality eligible, retained complete trajectories, passed separate hidden verification, produced no credential/canary leakage, and confirmed sandbox cleanup. Mogil Bench v0.1.0 was published as a verified GitHub public-alpha pre-release on 2026-07-15: https://github.com/mogilventures/mogil-bench/releases/tag/v0.1.0.

Non-goals

  • BlindBench executing agents or owning model/runtime credentials.
  • Training or deploying a model without explicit approval.
  • Customer/private traces before data-boundary approval.
  • Rebuilding Harbor environment orchestration inside Mogil Bench.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions