Labelbox · Engineering profile

Judge agents with data you can trust.

Mohamed A M Elansary, PhD — production agent evaluation sets, scientific/HPC data pipelines, and uncertainty-aware failure analysis.

Agent evaluation setsScientific data pipelinesUncertainty quantificationPython systems

Production agent evals

  • Builds GPT, Claude, and Gemini agent workflows at Vertexium.
  • Maintains regression evaluation sets for routing, retrieval, and isolation failures.
  • Ships provenance-aware Python pipelines with multi-tenant data isolation.

Scientific data infrastructure

  • Ran Linux/HPC pipelines over USGS, NOAA, and NASA datasets.
  • Quantified uncertainty and reported regime-dependent failure modes.
  • Validated imperfect observations before interpreting model performance.

Proposed eval approach

Define intended agent behavior and a failure taxonomy; run, store, and compare trajectories with provenance; attach uncertainty and task stratification; compare simple baselines; and report what the signal does and does not support before it is used as a product-quality or training input.

Honest fit boundary

Staff full-stack product engineering and TypeScript depth are a stretch. I have not built React/TypeScript product surfaces as a staff engineer, trajectory eval factories at Labelbox scale, or fine-tuning pipelines. I do not claim RLHF, safety research, or invented metrics. My contribution is production agent evaluation sets, scientific/HPC data pipelines, and failure analysis in Python.

Role and location

Staff Software Engineer, AI Data Platform · San Francisco Bay Area. Work mode is unstated in the posting. Relocation with a support package is an honest discussion point; remote or hybrid eligibility is not asserted.

“Annual base salary range $250,000 — $280,000 USD” · Official role posting