Project Category
LLM Evaluation & Agent Workflow Analysis
Engineering experience in Python-based task execution, LLM evaluation, reproducible benchmarking, and agent workflow failure analysis.
LLM Evaluation Workflows
Evaluation scope: task execution, instruction following, and prompt reasoning.
Work: evaluation workflows and reproducible benchmarking for LLM-driven tasks.
Agent Workflow Analysis
Evaluation scope: agent task completion and workflow reliability.
Work: rubric-based assessment, analysis of failure modes, and benchmarking of automated workflows.