Ritesh Srivastava
Machine Learning Engineer | LLM Evaluation & Applied ML Systems
Summary
Machine Learning Engineer with 7+ years building production LLM evaluation systems, model-quality scoring, and automated ML pipelines. At Adobe, lead development of LLMJury and PRISM — evaluation frameworks that apply statistical thresholding, LLM-as-a-judge triage, and bias and attribution scoring to certify Generative AI features. Reduced evaluation cycles by 75–90% and cut confirmed regression-review volume by half across 13 locales.
Core Skills
LLM evaluationGenerative AI evaluationModel quality scoringLLM-as-a-judgeStatistical thresholdingBias & fairness evaluationAttribution & relevance scoringNLP
PythonPandasNumPySciPyScikit-learnPySparkSQL
Weights & BiasesVoxel51LabelboxPrefectDockerAWSGit
Experience
Adobe Systems — Document Cloud & Artificial IntelligenceNoida, India
Lead Software EngineerJul 2025 – Present
- Lead PRISM's evaluation-scoring engine, rolling per-metric delta thresholds into auditable GREEN/YELLOW/RED release decisions across 13 locales and 2 pipelines.
- Implemented LLM-based triage (Claude Sonnet) to classify flagged changes as regressions, improvements, or false alarms, halving confirmed regressions requiring review.
- Built attribution and relevance scoring with 86.6–94.6% coverage, plus bias evaluation across 13 locales, extending evaluation from accuracy to fairness and grounding.
- Defined a reusable evaluation dataset strategy — blindset, introset, realworldset, and biasset — adopted across certified GenAI verticals.
- Automated per-locale reporting, replacing person-weeks of manual Excel scoring and reducing evaluation time 75–90% (1–2 days to 3–4 hours).
Senior Software EngineerApr 2022 – Jul 2025
- Led development of LLMJury, a multi-model evaluation framework for Generative AI summary, podcast, and assistant features.
- Improved the QA pipeline to run 30× faster and achieve 15% greater precision than human evaluators, catching critical regressions before release.
- Built a Voxel51/Labelbox-integrated analysis tool that cut manual evaluation effort 80%.
- Applied core ML/NLP techniques to Acrobat's AI features, supporting model analysis and data-driven decisions.
Tata Consultancy ServicesNoida, India
System Engineer — Walgreens, Corteva/DuPont, EatonNov 2018 – Apr 2022
- Built Python automation and data-processing scripts for enterprise clients, including EDA workflows, PySpark batch jobs, and Django REST APIs.
Selected Projects
- PRISM — agentic evaluation scoring and triage engine; LLM-as-a-judge, statistical thresholds, attribution metrics, and release decisions.
- LLMJury — multi-model Generative AI evaluation framework; 30× faster and 15% more precise than human review.
- Voxel51 & Labelbox annotation tooling — ML evaluation and dataset annotation pipelines.
- Kaggle EDAs & PySpark jobs — applied data analysis and distributed data processing.
Education
B.Tech, Computer Science & Engineering — FGIET, AKTU, Uttar Pradesh2014 – 2018
Certifications
Certified MLOps Practitioner — Adobe Cohort 01, School of DevOps
Python & Django Full Stack Web Developer — Udemy