Senior Evaluation Algorithm Engineer
Hong KongRemote
AI Summary
Build LLM evaluation systems for dialogue and financial trading scenarios at Binance, designing metrics, rubrics, datasets, and automated evaluation workflows to drive model improvement.
About this role
Binance is a leading global blockchain ecosystem behind the world’s largest cryptocurrency exchange by trading volume and registered users. We are trusted by 300+ million people in 100+ countries for our industry-leading security, user fund transparency, trading engine speed, deep liquidity, and an unmatched portfolio of digital-asset products. Binance offerings range from trading and finance to education, research, payments, institutional services, Web3 features, and more. We leverage the power of digital assets and blockchain to build an inclusive financial ecosystem to advance the freedom of money and improve financial access for people around the world.
About the Role
In the AI era, large language models are reshaping core business scenarios such as dialogue and trading. Model capability iteration relies on a scientific and trustworthy evaluation system — the "ruler" that measures model quality and guides R&D direction. We are seeking an evaluation expert with an algorithmic background to build LLM evaluation capabilities covering dialogue, financial trading, and other scenarios, using professional evaluation methods to quantify model performance, pinpoint issues, and drive continuous model improvement.
Responsibilities
- Design end-to-end LLM evaluation plans for business scenarios such as dialogue and financial trading. Build evaluation metric systems and rubrics, transforming subjective model performance judgments into quantifiable, reproducible, and explainable evaluation conclusions.
- Lead the design and construction of evaluation datasets. Define evaluation dimensions and scenario coverage, establish high-quality data annotation guidelines and quality control processes, and build benchmarks that authentically reflect business needs and have discriminative power.
- Analyze model capability boundaries and failure modes based on evaluation results. Produce actionable improvement recommendations and collaborate with algorithm and product teams to drive model iteration, making evaluation a critical component of the R&D loop.
- Drive the automation and scaling of evaluation workflows. Build sustainable evaluation platforms and toolchains to support high-frequency, stable evaluation needs during rapid model iteration.
- Collaborate with algorithm, product, and data teams to translate business and model objectives into clear evaluation standards, and turn evaluation findings into concrete R&D directions and drive their implementation.
Requirements
- Master's degree or above in Computer Science, Artificial Intelligence, Mathematics, Statistics, or related fields, with a solid algorithmic foundation and understanding of LLM principles, training, and fine-tuning processes.
- Hands-on LLM evaluation experience at a large tech company, with participation in commercial deployment evaluation (not purely academic or offline benchmarking). Familiar with the full pipeline from evaluation data preparation and rubrics design to evaluation-driven R&D.
- Familiar with mainstream evaluation methods (human evaluation, model-based automatic evaluation / LLM-as-a-judge, metric computation) and their applicable boundaries. Able to define appropriate evaluation dimensions for different business scenarios and write clear, actionable, and discriminative rubrics.
- Systematic control over evaluation data representativeness, annotation consistency, and result reliability, ensuring scientific and trustworthy evaluation conclusions.
- Proficient in Python, with experience in evaluation workflow automation, benchmark construction, or evaluation platform development. Able to independently handle data processing, evaluation script writing, and result analysis.
- Strong business understanding and communication skills, able to translate evaluation findings into clear improvement directions and effectively drive cross-team collaboration.
Bonus Qualifications
- Experience evaluating dialogue systems, AI Agents, or financial/trading LLMs.
- Experience building high-quality AI training/evaluation data or data annotation systems.
- Familiarity with RLHF, reward models, or preference data-related work.
Skills
Annotation Quality ControlBenchmark ConstructionData AnnotationData ProcessingEvaluation Metric SystemsEvaluation Platform DevelopmentEvaluation Workflow AutomationHuman EvaluationLLM-as-a-JudgeLLM EvaluationMetric ComputationModel-based Automatic EvaluationPythonReward ModelsRLHFRubrics Design
Explore related jobs
More jobs at Binance
Similar Annotation Quality Control jobs
Jobs in Hong Kong
VFX and AI artistidNerd Studio Ltd. · Lai Chi Kok, Kowloon- Vice President, Regulatory ComplianceUnited Overseas Bank (Malaysia) Bhd · Hong Kong (City Area)
Assistant Manager / Manager, Data ScientistWeLab · Quarry Bay, Hong Kong
Senior UI/UX DesignerMox Bank · Hong Kong (SAR)- Senior Associate, PayrollWPP Media · Hong Kong, Hong Kong
- Senior Staff, Platform Engineer, DevOpsMoneyHero Group · Hong Kong
Browse these categories
Market data for this role
All reports →- SeriesRole reportsOne role family at a time: how many openings, what changed this week, who is hiring, what it pays.
- SeriesSalary reportsWhat employers publish in job postings, by level and workplace. Not self-reported pay.
- Market overviewState of tech hiring, September 2026: up 4.8%Tech hiring rose 4.8% month over month in September 2026, with 411,122 new listings. Customer support and account executive roles led the growth.
