Sezer Mikail SagirBruchsal, Germany

  • LLM quality review
  • Annotation QA
  • DE · TR · EN · JA

AI evaluation that catches what automated metrics miss.

I find the errors your metrics wave through. You get a clear report with failure patterns and concrete fixes.

  • Fixed price to start
  • NDA, no problem
  • Quote within 1–2 days
Portrait of Sezer Mikail Sagir
Sezer Mikail SagirBruchsal, Germany

15–25%

Fewer annotator errors

~20%

More consistent grading

3,000+

Structured evaluations

Approximate figures from freelance AI evaluation work via RWS and Volga Partners, April 2025 to present. Client and program details are under NDA.

Worked through

RWS Volga Partners Scale AI micro1 OneForma Agents Only

Sound familiar?

The most expensive failures in an AI product are the ones that look green on the dashboard.

Hands on a laptop showing source code, reading glasses beside it

The metric says “pass”. Users still complain.

Scores miss dropped constraints and invented details.

Your annotation team grades inconsistently.

Vague guidelines produce training data nobody trusts.

Nobody checks the German, Turkish or Japanese outputs.

Unnatural translations simply slip through.

That is exactly where I come in.

What you get.

Direct, no agency. You talk to the person who reads your data.

AI evaluation

Every quality axis scored separately, with named failure patterns and feedback your team can retrain on.

3,000+ evaluations in production RLHF programs

Annotation QA & guidelines

Sample reviews, per-annotator feedback and clear rules for recurring disputes.

15–25% fewer errors in teams I coached

Review tools & LLM workflows

Fast review interfaces and LLM steps with built-in checks, down to the infrastructure.

Example: the TTS Review Tool

The low-risk way to start

Start with a pilot audit.

See what a close review finds in your own data before you commit to more.

What’s included

  • Review of 100 outputs from your model, per quality axis
  • A report with the most frequent failure patterns and real examples from your data
  • Concrete suggestions for prompts, guidelines or training data
  • A short call about the results, if you’d like one
Hand writing notes in a notebook with a pen

from €590

Fixed price · report within 5 business days

Request a pilot audit

Afterwards, ongoing weekly or monthly if you want. Larger scopes: quote within 1–2 days.

How it works.

  1. 1

    Request

    5 minutes, via form or email

  2. 2

    Quote

    within 1–2 days, NDA up front if needed

  3. 3

    Pilot audit

    100 outputs in 5 business days

  4. 4

    Report

    Failure patterns, examples, fixes

Evidence, not promises.

Client work is under NDA. Here is the format of my reports and a tool from my daily work.

Evaluation case studiesNDA-safe

Two outputs that look the same – until you read them.

Six failure patterns and two reviews, line by line in the format of my real reports.

Evaluation toolingSep 2026

TTS Review Tool

  • Live demo
  • Open source
  • In daily use

Keyboard-driven A/B rating for text-to-speech: play, score, tag, next. Without touching the mouse.

CMOS, A vs. Bkeys 1–5

Keymapfrom the README

  • Play A → BSpace
  • Play only A / only BAB
  • Set CMOS score1–5
  • Toggle reason tagsQ…]
  • Previous / next item←→
  • Confirm comment & advanceEnter

Schematic of the review loop. Try the real interface in the live demo.

Who reviews your data.

Since April 2025 I have reviewed at senior level on large generative AI evaluation programs. My focus: the gap between “scores well” and “is actually right”. Recurring failures become clear guidelines.

Résumé (PDF)

Based in
Bruchsal, Germany. Remote, CET.
Focus
RLHF evaluation, rubric review, error analysis, guidelines
Platforms
SageMaker Ground Truth, Scale AI, Labelbox, SuperAnnotate
Languages
German, Turkish, English (C1), French (B2), Japanese (N3)

FAQ.

What does it cost?

The pilot audit is a fixed price from €590 for 100 outputs. For ongoing work you get a quote within 1–2 days.

Do you sign NDAs?

Yes, happily before the first data point. I have worked under NDA continuously since 2025.

Can you work in our tools?

Yes, e.g. SageMaker Ground Truth, Scale AI, Labelbox or SuperAnnotate. A spreadsheet works too.

Who is this for?

AI teams training models, and companies building LLM features that need to prove the quality holds up.

Request a pilot audit or discuss a project.

Briefly describe the task, language and volume. You get a reply with a quote within one to two days.

  • Fixed price for the pilot
  • NDA available
  • Direct contact, no agency
  • Remote, Central European Time

Prefer a call? Mention a few times that work for you and I’ll send a scheduling link.

Send a message

A few lines on the project, the scope and the timeline are enough.

Your details are only used to reply to your message. See the privacy policy.