Board
GEPA Prompt Lab
GEPA Prompt Lab
Paste a prompt. Get a better one.
✎
Improve a Prompt
Your Prompt
Feedback
(optional)
Improve
✓
Improved Prompt
Copy
Iterate
Show Advanced Mode ►
✎
Prompt Editor
i
Your starting prompt. GEPA generates mutations, evaluates each candidate against your training examples, selects the best, and iterates — returning an improved prompt.
-- Load from library --
Delete
Name
i
A short label for this prompt in your library. Pick something descriptive — you'll select it by name when setting up a run.
Seed Prompt
i
Your initial prompt draft — it doesn't need to be perfect. GEPA will optimize it. Write something that roughly describes the task (e.g. "Summarize the following article in 3 bullet points.").
Save to Library
☷
Training Data
i
Input/output pairs that teach the evaluator what a good response looks like. At least 5 examples recommended. GEPA uses 80% for training and 20% for validation.
-- Load set --
New Set
Delete Set
Set Name
Input
i
The text or task you'd give to the LLM in production (e.g. an article to summarize, a question to answer). GEPA runs each prompt candidate against these.
Expected Output
i
An example of what a good response looks like. The LLM judge compares candidates against this reference when scoring each eval dimension.
+ Add Example
Save Set
✓
Eval Config
i
Defines how prompt candidates are scored. Each dimension is a criterion (e.g. Clarity, Accuracy) rated 0–10 by an LLM judge. GEPA optimizes toward the weighted composite score.
-- Load preset --
General Prompt Quality
Image Prompt
Summary Quality
-- Load saved --
Delete
Config Name
i
A name for this evaluation setup. Saved to your library so you can reuse it across multiple optimization runs.
+ Add Dimension
Save Config
⚙
Model & Settings
i
Configure which models GEPA uses and how much compute budget to allocate. Higher metric call budgets produce better results but cost more and run slower.
feat-005
Task Model
i
The model that runs each prompt candidate against your training examples. Format: provider/model-name (e.g. azure/gpt-4o-mini, openai/gpt-4o).
Reflection Model
i
The model used as LLM judge to score each candidate's outputs against your eval dimensions. Can be the same as the task model.
Max Metric Calls —
50 calls
i
Total LLM scoring calls during the optimization run. Higher = more thorough search, better results, but slower and more expensive. 50 is a good starting point; use 100+ for production prompts.
Selection Strategy
i
Pareto
: selects prompts that score well across all eval dimensions simultaneously — best when you want balanced quality.
Lexicase
: selects prompts that excel on specific examples — good when training data is varied and you want specialist prompts.
pareto
top-k pareto
▶
Run Controls
i
Once Prompt, Training Data, and Eval Config are saved, click Optimize. GEPA runs a genetic loop: generate candidates → evaluate → select best → mutate → repeat. Results appear below when done.
Run Name
i
A label for this optimization run. Saved to Run History so you can compare results across iterations and track which settings produced the best prompts.
Stop
Optimize
★
Results
Copy Optimized
Save to Library
Seed Prompt
Optimized Prompt
🕑
Run History
Refresh
Name
Date
Score
Status
No runs yet.