Model Evaluation Engineer
You will develop evaluations that make model quality, regressions, and tradeoffs easier to measure.
About the role
The work includes evaluation design, datasets, scoring, execution tools, dashboards, anomaly investigation, and clear communication of results.
What you'll do
- Define evaluation tasks, representative datasets, scoring methods, and validation checks.
- Develop reliable tools for running evaluations and comparing results.
- Investigate anomalous results and determine whether the cause is model behavior, data, or evaluation software.
- Create reports and visualizations that make evaluation results useful to researchers and decision-makers.
What we're looking for
- Strong Python skills and experience developing research or production tools.
- Experience with evaluation design, experimental methods, or quality measurement.
- Strong attention to detail and the ability to turn a fuzzy quality judgment into measurable criteria.
- Clear written and verbal communication about technical results.
Preferred experience
- Experience with statistics, data visualization, monitoring, or experiment tracking.
Equal opportunity
Entlegnant, the company behind Leiolai, is an equal opportunity employer. We consider qualified applicants without regard to race, color, religion, sex, gender identity or expression, sexual orientation, national origin, ancestry, age, disability, veteran status, genetic information, or any other status protected by law.
Applicants must be at least 18. If you need an accommodation during the application process, email talent@leiolai.com.
Apply
Submit your resume and a few details below. A cover letter is optional.