PUT PROMPTS TO THE TEST
MVP in developmentAI evaluation and developer tools
AI Judge
Measure a prompt by what its code can do.
An evaluation platform that tests AI-generated code against problem cases and compares submissions and rankings.
HOW IT COMES TOGETHER
The core workflow- 01
Problem and prompt
- 02
LLM-generated code
- 03
Test-case evaluation
- 04
Records and comparison
01 / THE QUESTION
The problem we set out to solve.
Plausible-looking code may not solve the actual problem. Comparing prompts or models requires shared problems and consistent evaluation criteria.
02 / OUR APPROACH
Our approach.
A submitted prompt asks an LLM to generate code, and a judge runs the test cases. This MVP records the path from generation to verdict so results can be compared.
- 01
Problems and prompt submissions
Problem listings, details, and submission screens guide users into an evaluation.
- 02
Model provider connections
Mock responses and an OpenAI-compatible provider separate local development from real model calls.
- 03
Test-based judging
Generated code is checked by a local judge, with an integration path for an external HTTP judging service.
- 04
History and comparison
Submission records, model and ranking screens, and the foundations for contests and administration make results easier to compare.
03 / INTELLIGENCE AT WORK
Where AI does the work.
The code produced by AI is the subject of the evaluation. The core idea is to compare model and prompt outputs against executable tests.
04 / WHERE WE ARE
Where the project stands.
An MVP connecting problem discovery, submissions, judging, and comparison. The default environment uses a mock LLM and local judging. Execution isolation and security validation remain necessary before a public launch.
