Skip to main content
AI Agents & Automation

Prompt Engineering andEvaluation

Replace trial-and-error prompting with versioned prompts and a test suite, so every change is measured and regressions are caught before users see them.

Prompt Engineering and Evaluation: what the work involves

Many LLM features are held together by a prompt that one person edited at midnight. It works on the examples they tried, behaves oddly on others, and nobody can say whether the last tweak helped. A model upgrade or a new type of customer message quietly changes the output, and the team finds out from complaints. Without measurement, every change is a gamble and every fix risks another break.

DevKey brings test discipline to prompts. We define what good output means for your task, collect real inputs and write expected behaviours, then build an evaluation set mixing exact checks, rubric scoring by a model and human review for the subtle cases. Prompts are broken into versioned components with structured outputs and guardrails. Each change runs against the suite in your pipeline, with results compared against the previous version, and we hand over the harness and the habits to maintain it.

What we build

Core features

01

Task definition and rubrics

We turn vague goals like helpful and accurate into written criteria and examples that people and automated judges can apply consistently.

02

Representative test set

Real inputs, including difficult and adversarial ones, are collected and labelled, and a portion is held back to avoid tuning to the test.

03

Structured prompt design

Instructions, examples and output schemas are separated and versioned, and outputs are validated against formats before use.

04

Automated evaluation harness

Exact-match, schema, rule and model-graded checks run on every prompt or model change, producing a comparison report.

05

Guardrails and refusal behaviour

We test how the system handles out-of-scope requests, injected instructions and sensitive topics, and tighten behaviour where it fails.

06

Model comparison and cost analysis

The same suite is run across candidate models, showing quality, latency and cost per request so you can pick on evidence.

Planned for

What we get right before launch

Model judges are imperfect

Using one model to grade another is fast but biased and sometimes wrong. We calibrate judges against human labels on a sample, report agreement, and keep periodic human review in the loop.

Overfitting to the test set

Tuning a prompt until it passes the same examples proves little. We keep a hidden holdout, refresh the suite with production failures, and watch for scores that rise without real improvement.

Evaluation data may be sensitive

Test sets drawn from real conversations contain personal data. We redact or synthesise where possible, store the suite with proper access control, and agree retention rules.

Stack

Tools and technology

  • Python
  • OpenAI GPT
  • Anthropic Claude
  • Google Gemini
  • Promptfoo
  • LangSmith and Langfuse
  • pytest
  • MLflow
  • GitHub Actions
Prompt Engineering and Evaluation FAQ

Common questions, answered

Is prompt engineering still relevant with newer models?

Yes, though the emphasis has shifted from clever phrasing to clear specification, good examples, structured outputs and measurement. Newer models forgive sloppy prompts more, but still need testing for your particular cases.

How large does an evaluation set need to be?

It depends on variety and risk. A few dozen well-chosen cases reveal many problems, while high-stakes tasks need hundreds. We start small, grow it from production failures, and avoid claiming precision a small set cannot support.

Can you evaluate Urdu and Roman Urdu outputs?

Yes, using native-speaking reviewers for the human portion. Automated judges are less reliable in Urdu, so we calibrate them against human scores before trusting them for regression checks.

Will you work with our existing system?

Yes. We can add the evaluation harness around prompts already in production, starting by capturing real inputs and outputs, then writing tests that protect the behaviours you cannot afford to lose.

Ready to start your Prompt Engineering and Evaluation project?

Tell us what you need and we will come back with a clear scope, timeline and the questions worth answering before any build starts.