Experience
Where I've been.
Work
- Jul 2026 – Present
Al Evaluation & LLM Benchmarking Specialist.
Handshake AI (Freelance, Remote) · July 2026 – PresentWorked as a Freelance AI Evaluation Engineer on Handshake AI's Project Dynamo, designing software engineering benchmarks for coding-focused large language models. Built deterministic evaluation pipelines, calibrated benchmark difficulty, and validated Docker-based benchmark environments to support reliable model evaluation.
- Authored software engineering benchmark tasks covering debugging, feature implementation, repository maintenance, and code understanding.
- Developed deterministic pytest verifiers and reference implementations for automated evaluation.
- Calibrated benchmark difficulty against reference coding agents to ensure rigorous and reproducible evaluation.
- Built and validated self-contained Docker environments for reproducible benchmark execution.
- Performed end-to-end validation using Python, Git, Linux, Docker, and CI workflows before submission.
- Evaluated and benchmarked coding-focused large language models to improve reasoning, debugging, and code generation capabilities.
PythonGitGitHubLinuxDockerPytestCI/CDBashJSONYAMLMarkdownSoftware DebuggingBenchmark EngineeringAI EvaluationLarge Language Models (LLMs)Terminal-Based Development - Aug 2024 – Sep 2024
Data Science and Machine Learning Intern
YBI Foundation · RemoteCompleted a Data Science and Machine Learning internship, building a Movie Recommendation System end-to-end.
- Built a hybrid recommender using collaborative and content-based filtering.
- Preprocessed data and engineered features for movie metadata.
- Explored NLP techniques for computing similarity across titles and descriptions.
- Documented methodology, evaluation metrics and results.
PythonPandasNumPyScikit-learnNLP