Roberto Aguirre Guardia
GenAI Quality & Evaluation Specialist · Testing LLM & Agent Outputs · Hallucination, Regression & Guardrail Checks · Human in the Loop · Langfuse Observability & Evaluation · Deep LLM / RAG Understanding · Systems Engineer + MBA
Bogotá, Colombia · Remote / Freelance aguirrerjg@gmail.com +57 317 675 0419 linkedin.com/in/roberto-javier-aguirre-guardia

Summary

GenAI quality and evaluation specialist. I build AI systems and own their quality: I evaluate and validate LLM and agent outputs, design checks for hallucinations, wrong outputs and regressions, and build human in the loop validation. I coordinate autonomous QA agents (Agent Squad) in production, with evaluation and traceability in Langfuse and MLOps (model versioning, agent monitoring, CI/CD). I am deep on LLM behavior (prompt engineering, RAG, function calling, embeddings), so I know where GenAI breaks and how to test for it: edge cases, adversarial inputs and guardrail failures. 15+ years delivering software to production with a high quality bar. Systems Engineer with an MBA. Available for freelance engagements.

Experience

Head of AI · GenAI Quality, Evaluation & QA in Production · DigitalHubAssist LLC Oct 2024 · Present
Remote · AI / SaaS (Anthropic partner program)
Consulting Manager · Software Delivery Quality & Defect Management · NTT DATA Colombia Apr 2021 · Oct 2024
Bogotá, Colombia · Consultancy (GSI) · Client: banking (Grupo AVAL)
Project Manager / Digital Transformation Lead · Testing & Go Live to Production · VASS (2018·21) · GlobalHitss (2015·18) · Tecnocom (2012·15) 2012 · 2021
Bogotá, Colombia · Pensions · Health · Public Sector

Education

MBA · Tecnológico de Monterrey · Systems Engineering, U. de Lima (Top Third) · Founding Professor of Innovation, U. Sergio Arboleda 2004 · 1999

Core Competencies

GenAI Output Evaluation (Hallucination · Regression · Correctness) QA of LLM & Agent Systems · Autonomous QA Agents (Agent Squad) Human in the Loop Validation · Structured & Auditable Outputs (JSON · Logs) Langfuse Observability & Evaluation · Model Versioning · Monitoring Adversarial & Edge Case Testing · Guardrail Checks LLM Behavior Depth (Prompt Engineering · RAG · Function Calling · Embeddings) Software Delivery Quality · Defect Management · Testing to Production Systems Engineer · MBA · Fluent English · Available Freelance