
Posted Today
Data Scientist– AI Infra & Evaluation Foundations
AI Summary
Designs evaluation frameworks and metrics for agentic AI systems, partners with engineering and product teams to standardize AI agent testing and monitoring, and builds end-to-end eval pipelines to ensure production AI is trustworthy and reliable.
About this role
About monday.com:
monday.com is the AI work platform powering the most ambitious teams. 250,000+ customers across departments use us to bring people, workflows, and AI agents together on one flexible platform where AI doesn't just assist, it executes. We move fast, build things that matter, and foster an ownership-driven culture where you're empowered to shape how organizations work and outpace their competition.
About the team
The AI Infra group builds the foundations, tools, and platforms that every team at monday relies on to ship intelligent, agentic features. We own the core infrastructure—including the AI Gateway and our centralized Evals framework—ensuring every AI feature deployed to production is secure, resilient, cost-effective, and above all, trustworthy.
Our focus is on the frontier of agentic AI: dissecting complex agent trajectories, building robust evaluation frameworks, and turning subjective notions of "good AI" into rigorous, actionable metrics. As a Data Scientist on this team, you will bridge the gap between AI research and production infrastructure. You'll partner closely with engineering and product teams across monday to design the judges, metrics, and error analysis workflows that allow us to ship cutting-edge AI agents with speed and confidence.
This position is based at our Tel Aviv office (Headquarters).
About the role:
As a Data Scientist in AI Infra, your goal goes far beyond simply building an evaluation framework—you will own the organizational impact of how monday evaluates and trusts AI. You will define how teams measure quality, influence engineering decisions across R&D, and turn fuzzy notions of "good AI" into numbers product teams rely on to ship with confidence.
Own the evaluation methodology: Design metrics, pipelines, and methodology that teams across monday trust and adopt as their source of truth.
Transform the AI agent lifecycle: Standardize how AI agents are built, regression-tested, and maintained across the org, embedding continuous evaluation into everyday engineering workflows and post-deployment monitoring.
Drive organizational impact & enablement: Partner with AI feature teams across monday to translate domain expectations into meaningful datasets, test suites, and continuous evaluation pipelines—leveling up engineers and product managers along the way.
Build hands-on tools: Prototype and stand up eval pipelines end-to-end, bridging the gap between ambiguous, high-level product requirements into clear, quantifiable evaluation standards that become central to how features are greenlit for production
Anticipate future failure modes: Stay ahead of evolving agent architectures by proactively designing next-generation evaluation strategies.
Requirements:
Agentic Systems & Architecture:
3+ years of experience as a Data Scientist in non-academic settings working with complex production running AI systems. Familiarity with current agentic frameworks like LangGraph, LangChain, and SoTA SDKs.
Deep, practical understanding of how agents operate - models, context, capabilities and harnesses. Deep experience with agentic evaluation methodologies
Execution, Code & Trace-First Mindset:
Production-grade coding skills with a track record of building, prototyping, and shipping end-to-end data or eval pipelines,
A "trace-first" diagnostic mindset—comfortable diving into raw agent execution logs, inspecting failure modes, and constructing qualitative error taxonomies.
Product & Organizational Impact:
Strong product intuition and exceptional communication skills to translate complex evaluation data into clear, actionable guidelines.
Proven ability to partner closely with software engineers and product teams, taking ownership of driving adoption and raising the quality bar across the organization.
Preferred Qualifications
Direct experience designing evaluation strategies for complex agentic systems in production
Prior experience working within centralized platform/infra teams that support multiple product verticals.
Experience contributing to modern microservice architectures, GitHub workflows, and automated production CI/CD pipelines
Familiarity with modern agent and eval tooling and observability stacks (e.g., LangSmith, Langfuse, or custom internal platforms).
Knowledge of TypeScript or experience working within modern platform architectures.
Master’s degree in Computer Science, Data Science, Statistics, Engineering, or a related quantitative field.
#LI-DNI
Skills
Explore related jobs
More jobs at monday.com
Similar Agentic Evaluation jobs
Jobs in Tel Aviv
- Senior Web Application Engineer - ILIbex Medical Analytics · Tel Aviv, Tel Aviv
- LCost Manager: (Electrical Focus)Linesight · Tel Aviv, Israel
- LSenior Cost Manager: (Electrical Focus)Linesight · Tel Aviv, Israel
- GProgram Assistant - NYU Tel AvivGlobal Nyu · Tel Aviv, Israel
- GWellness Counselor - NYU Tel AvivGlobal Nyu · Tel Aviv, Israel
- Software Engineer, DataMONTICELLOAM · Tel Aviv, Israel
Browse these categories
Market data for data scientist roles
All reports →- SeriesRole reportsOne role family at a time: how many openings, what changed this week, who is hiring, what it pays.
- SeriesSalary reportsWhat employers publish in job postings, by level and workplace. Not self-reported pay.
- Market overviewState of tech hiring, September 2026: up 4.8%Tech hiring rose 4.8% month over month in September 2026, with 411,122 new listings. Customer support and account executive roles led the growth.