Mehdi Akiki

Posts tagged "LLM evaluation"

  • Published on
    A measurement architecture that joins repeatable offline evals with production outcomes, safety signals, canaries, and privacy-aware review.
  • Published on
    A retrieval benchmark for separating missing evidence from generation failures using Recall@k, Precision@k, MRR, unanswerable questions, and controlled oracle context.
  • Published on
    A graded assertion approach for regression-testing variable AI outputs without comparing every generated word to one frozen answer.
  • Published on
    A working record-and-replay harness for agent tests: a JSON cassette of decisions and tool results, deterministic replay with zero live calls, and a readable diff when the trajectory changes.