Apple · AI Quality Engineer / SDET
I own the evaluation framework for an internal order-operations assistant used by support operations and QE. It scores six dimensions in Python, keeps five of them deterministic, runs a 600-item benchmark with derived golden answers, and gates the pipeline on the numbers that matter. Alongside it I maintain MCP server integrations for agent tooling and automate backend validation across 30+ service clients over gRPC and REST.