
Posted 3 months ago
Applied AI Software Engineer
San FranciscoHybridFull-time
AI Summary
Lead evaluations for AI agents in development and production at Canvas Medical’s EMR and payments platform. Design and run large-scale experiments measuring performance, safety, and reliability of LLM-based agents across clinical, operational, and financial workflows.
About this role
Canvas Medical is the electronic medical records (EMR) and payments development platform for healthcare. We build modern, elegant front- and back-end tooling to enable new ways for developers and clinicians to collaborate to solve healthcare’s toughest challenges. Canvas is institutionally backed by some of the greatest technology investors in the world (funded notable health tech companies such as GoodRx, Oscar Health, and Hims & Hers Health).
The Role
We’re hiring an Applied AI Software Engineer to lead evaluations for agents in development and the post-deployment fleet of agents operating in Canvas to automate work for our customers. You will help develop agents in Canvas using state of the art foundation model inference and fine-tuning APIs along with our server-side SDK. The server-side SDK provides extensive tools and virtually all the context necessary for excellent agent performance. You’ll be responsible for designing and running rigorous evaluation experiments that measure performance, safety, and reliability across a wide variety of clinical, operational, and financial use cases.
This role is ideal for someone with deep experience evaluating LLM-based agents at scale. You’ll create high-fidelity unit evals and end-to-end evaluations, define expert-determined ground truth outcomes, and manage iterations across model variants, prompts, tool use, and context window configurations. Your work will directly inform model selection, fine-tuning, and go/no-go decisions for AI features used in production settings.
You’ll collaborate with product, ML engineering, and clinical informatics teams to ensure that Canvas's AI agents are not only capable, but trustworthy and robust under real-world healthcare constraints. You will also work with technical product marketers and developer advocates to help our broader developer community and the broader market understand the uniquely differentiated value of agents in Canvas.
Who You Are
What You’ll Do
What Success Looks Like at 90 Days
Qualifications
Skills
Experiment TrackingFoundation ModelsHealthcare DataLLMLLM EvaluationModel Fine-tuningPrompt EngineeringPythonReinforcement LearningSQL
Explore related jobs
More jobs at Canvas Medical
Similar Experiment Tracking jobs
Jobs in San Francisco
Proposal ManagerPGH Wong Engineering, Inc. · San Francisco, California
Software Engineer Intern, Mobile (Winter 2027)Notion · San Francisco, California- Residential Security Agent (San Francisco, CA)Concentric · San Francisco, California
Partner Marketing Manger, GSI & SIAnthropic · San Francisco, California | New York City- Marketing ManagerAsset Living · Denver, CO
Program Manager, ConnectMeter · San Francisco
Browse these categories
Market data for ai / ml engineer roles
All reports →- SeriesRole reportsOne role family at a time: how many openings, what changed this week, who is hiring, what it pays.
- SeriesSalary reportsWhat employers publish in job postings, by level and workplace. Not self-reported pay.
- Market overviewState of tech hiring, September 2026: up 4.8%Tech hiring rose 4.8% month over month in September 2026, with 411,122 new listings. Customer support and account executive roles led the growth.