Judge agents with data you can trust.
Mohamed A M Elansary, PhD — production agent evaluation sets, scientific/HPC data pipelines, and uncertainty-aware failure analysis.
Production agent evals
- Builds GPT, Claude, and Gemini agent workflows at Vertexium.
- Maintains regression evaluation sets for routing, retrieval, and isolation failures.
- Ships provenance-aware Python pipelines with multi-tenant data isolation.
Scientific data infrastructure
- Ran Linux/HPC pipelines over USGS, NOAA, and NASA datasets.
- Quantified uncertainty and reported regime-dependent failure modes.
- Validated imperfect observations before interpreting model performance.
Proposed eval approach
Define intended agent behavior and a failure taxonomy; run, store, and compare trajectories with provenance; attach uncertainty and task stratification; compare simple baselines; and report what the signal does and does not support before it is used as a product-quality or training input.
Honest fit boundary
Staff full-stack product engineering and TypeScript depth are a stretch. I have not built React/TypeScript product surfaces as a staff engineer, trajectory eval factories at Labelbox scale, or fine-tuning pipelines. I do not claim RLHF, safety research, or invented metrics. My contribution is production agent evaluation sets, scientific/HPC data pipelines, and failure analysis in Python.
Role and location
Staff Software Engineer, AI Data Platform · San Francisco Bay Area. Work mode is unstated in the posting. Relocation with a support package is an honest discussion point; remote or hybrid eligibility is not asserted.
“Annual base salary range $250,000 — $280,000 USD” · Official role posting