
Posted 24 days ago
AI Security & Control Researcher
AI Summary
Apollo Research seeks a security and control expert to design threat models and control protocols for coding agents, improve the Watcher monitoring product, and serve as a security adjudicator for flagged agent behavior.
About this role
THE OPPORTUNITY
THE OPPORTUNITY
Apollo Research works with most frontier AI companies (OpenAI, Anthropic, Google, Meta, Thinking Machines and others) to test their models before deployment and collaborate on fundamental scheming research. Our coding agent security product, Watcher, is deployed in production and monitors billions of agent tokens per month across engineering teams at agent-building scale-ups and enterprises.
We are looking for a security & control expert to help us design better threat models and control protocols against AI adversaries, and improve the effectiveness and security of Watcher.
This is truly a "start-up role" in the sense that you have significant say in shaping the direction of the role. This is an individual contributor role but could lead to management responsibilities eventually, if desired.
KEY RESPONSIBILITIES
Threat modeling for coding agents
Develop threat models for coding agents under the assumption that they could be misaligned or compromised. This includes near-term threats like AI agents leaking private internal data and long-term threats like agents sabotaging safety research.
Treat coding agents through the lens of insider risk: they have credentials, access to code, network access, and the ability to execute arbitrary actions, just like a malicious insider would (see Control agenda).
Map out kill chains and attack progressions similar to frameworks like MITRE ATT&CK, adapted for agentic AI. See e.g. the Agentic Loss-of-Control Threat Matrix for an example of a high-quality contribution.
Build and maintain our “coding agent security levels” which define what level of robustness Watcher provides against different categories of failure modes. These levels should be concrete, testable, and usable both internally (to guide product priorities) and externally (to communicate our security posture to customers).
Failure mode library and literature
Maintain our library of coding agent failure modes and keep it comprehensive, accurate, and current.
Stay up to date with the academic and industry literature on coding agent security, AI security, and adjacent fields (prompt injection, deception, accidental security implications, supply chain attacks on AI systems, etc.).
Attack design for monitor development
Design realistic attack trajectories and example exploits for monitor development and backtesting, work with the research team to ensure monitors are trained and evaluated against credible adversarial behavior.
Trajectory adjudication
Serve as security adjudicator for flagged trajectories. When monitors flag ambiguous agent behavior, determine whether it represents a genuine attack, a benign-but-unusual workflow, or needs deeper investigation, and feed those judgments back into monitor improvement.
Red-teaming Watcher
Red-team and improve Watcher's monitors and policies, document findings, and feed them into product improvement.
Work with the technical research staff to build adversarial test suites into the backtesting pipeline.
*This role is distinct from the AI Security Researcher role on Apollo's Infra & Security team, which threat-models Apollo's own internal systems and the agents operating on them.
REPRESENTATIVE PROJECTS
Develop a comprehensive coding agent threat model: Think deeply about all the ways in which coding agents could attack an organization. Compare this threat model to conventional threats from human insiders. Publish a detailed research piece describing the threat model building on existing research, e.g. from Redwood Research.
Improve our database of failure modes: We have an internal database of 50+ failure modes of coding agents with detailed reports for all of them. For this project, you would provide an expert view on the current state of that database and suggest improvements. In the long run, you would maintain that database and be responsible for integration of new failures.
Prioritize failure modes that Watcher should cover: Different parts of Watcher attempt to cover different threat models and attack strategies. Based on the results of the threat model project above, we want to ensure that each part of Watcher covers the most important failure modes in the most efficient way. For example, not all monitors require blocking capability and some failure modes might benefit from additional affordances like being able to disperse subagents.
JOB REQUIREMENTS
Must-haves
Nice-to-haves
Experience with AI/ML systems security, LLM security, or AI control research. The field is young enough that deep experience here is rare, but any exposure significantly reduces ramp-up time.
Detection engineering, SOC, or incident analysis experience. A part of this role is judging whether flagged agent behavior is genuinely malicious, and people who have triaged real-world alerts might ramp much faster.
Familiarity with insider threat programs or insider risk frameworks. The mental model of "the coding agent is a potentially malicious insider" is useful for this role and someone who has worked on insider threats will pick it up faster.
Red teaming or offensive security background. Useful for the Watcher red-teaming responsibilities and for thinking adversarially about failure modes.
Formal AI safety research background. Helpful but not necessary. We need security practitioners who can learn the AI safety context, not AI safety researchers who need to learn security.
Explicitly not required
Management experience. This is an IC role, at least initially.
Specific certifications (CISSP, etc.). We care about demonstrated ability, not credentials.
BENEFITS
LOGISTICS
Skills
Explore related jobs
More jobs at Apollo Research
Similar Adversarial Testing jobs
Jobs in London
- Enterprise Account Executive, French-Speaking - London, UKPallet · London, UK
- Enterprise Account Executive - Spanish-Speaking - London, UKPallet · London, UK
- Enterprise Account Executive, Portuguese-Speaking - London, UKPallet · London, UK
- Enterprise Account Executive - German-Speaking - London, UKPallet · London, UK
- Enterprise Account Executive - London, UKPallet · London, UK
Service Support AdministratorRentokil Initial Group · Woodford, London
Browse these categories
Market data for this role
All reports →- SeriesRole reportsOne role family at a time: how many openings, what changed this week, who is hiring, what it pays.
- SeriesSalary reportsWhat employers publish in job postings, by level and workplace. Not self-reported pay.
- Market overviewState of tech hiring, September 2026: up 4.8%Tech hiring rose 4.8% month over month in September 2026, with 411,122 new listings. Customer support and account executive roles led the growth.