PUT PROMPTS TO THE TEST

MVP in development

AI evaluation and developer tools

AI Judge

Measure a prompt by what its code can do.

An evaluation platform that tests AI-generated code against problem cases and compares submissions and rankings.

HOW IT COMES TOGETHER

The core workflow
  1. 01

    Problem and prompt

  2. 02

    LLM-generated code

  3. 03

    Test-case evaluation

  4. 04

    Records and comparison

01 / THE QUESTION

The problem we set out to solve.

Plausible-looking code may not solve the actual problem. Comparing prompts or models requires shared problems and consistent evaluation criteria.

02 / OUR APPROACH

Our approach.

A submitted prompt asks an LLM to generate code, and a judge runs the test cases. This MVP records the path from generation to verdict so results can be compared.

  1. 01

    Problems and prompt submissions

    Problem listings, details, and submission screens guide users into an evaluation.

  2. 02

    Model provider connections

    Mock responses and an OpenAI-compatible provider separate local development from real model calls.

  3. 03

    Test-based judging

    Generated code is checked by a local judge, with an integration path for an external HTTP judging service.

  4. 04

    History and comparison

    Submission records, model and ranking screens, and the foundations for contests and administration make results easier to compare.

03 / INTELLIGENCE AT WORK

Where AI does the work.

The code produced by AI is the subject of the evaluation. The core idea is to compare model and prompt outputs against executable tests.

04 / WHERE WE ARE

Where the project stands.

An MVP connecting problem discovery, submissions, judging, and comparison. The default environment uses a mock LLM and local judging. Execution isolation and security validation remain necessary before a public launch.