Project Category

LLM Evaluation & Agent Workflow Analysis

Engineering experience in Python-based task execution, LLM evaluation, reproducible benchmarking, and agent workflow failure analysis.

All projects
LLM AI Agents Prompt Engineering Evaluation Automation Reasoning

LLM Evaluation Workflows

Evaluation scope: task execution, instruction following, and prompt reasoning.

Work: evaluation workflows and reproducible benchmarking for LLM-driven tasks.

LLM Evaluation Prompt Engineering Reasoning

Agent Workflow Analysis

Evaluation scope: agent task completion and workflow reliability.

Work: rubric-based assessment, analysis of failure modes, and benchmarking of automated workflows.

AI Agents Automation Benchmarking Workflow Analysis