← All jobs
Appier

Senior Machine Learning Scientist (LLM & Agents)

Taipei, Taiwan
Salary
Not stated
Level
Senior
Work type
Not stated
Visa
Not stated

Open. First seen 7 October 2026.

About the role

About Appier

Appier is a software-as-a-service (SaaS) company that uses artificial intelligence (AI) to power business decision-making. Founded in 2012 with a vision of democratizing AI, Appier’s mission is turning AI into ROI by making software intelligent. Appier now has 17 offices across APAC, Europe and U.S., and is listed on the Tokyo Stock Exchange (Ticker number: 4180). Visit www.appier.com for more information.

About the Role

We are looking for a Machine Learning Scientist (LLM & Agents) to join the Enterprise Solution Science Team. This team focuses on applying cutting-edge ML technologies to real-world marketing problems by combining them with omnichannel customer data. In this role, you will own agent features end to end — from problem definition and design, to shipping, to measuring and improving quality in production.

What You’ll Work On

  • You will contribute to the design, development, and productionization of one or more enterprise AI agents, including:
  • Marketing Agent: Empower marketers to improve conversion rates and drive revenue growth.
  • Strategy Agent: Transform raw data into actionable, revenue-driving strategies and high-potential audience segments.
  • Data Analysis Agent: Answer business questions by planning multi-step analyses and running SQL over large customer datasets, with every number traceable to the query that produced it.
  • Sales and Service Agent: Drive sales conversion, increase average order value (AOV), and deepen brand engagement and customer loyalty through real-time, 24/7 AI-powered support.
  • Build agent capabilities as reusable skills and MCP tools, and design tool interfaces that agents use correctly and efficiently
  • Bring predictive models (e.g. propensity, churn, recommendation) into agent workflows and next-best-action suggestions
  • Build evaluation — test sets, LLM-as-a-judge, production trace review — and use it to decide what ships
  • Make agents faster and cheaper without losing answer quality, through context management, caching and model choice
  • Write design docs, break work into reviewable PRs, and review teammates’ code
  • Work closely with PMs, engineers and QA from planning to release

What We’re Looking For (Minimum Qualifications)

  • Bachelor’s degree or above in Computer Science, Machine Learning, Statistics, or a related field
  • 2+ years of experience in ML or software engineering roles
  • Has built at least one LLM or agent application end to end — at work, in research, or as a side project — and can explain how its quality was evaluated and improved
  • Hands-on experience building with MCP and agent skills (tools, instructions and reference files an agent loads on demand)
  • Hands-on experience with at least one agent harness (the runtime that drives the agent loop of model calls, tools and context), e.g. Claude Agent SDK, OpenAI Codex, or Google ADK — and an understanding of how the harness, tool design and context shape agent behavior
  • Can design evaluation for non-deterministic LLM output, including test sets, metrics, and the limits of LLM-as-a-judge
  • Strong SQL and data reasoning: can write and review analytical queries and tell when a result is wrong
  • Can profile and reduce end-to-end latency and cost of an agent — e.g. fewer model round-trips, better tool design, caching, model choice — without hurting answer quality
  • Writes production-quality Python with tests
  • AI-native development: works daily with coding agents that read the codebase, run tests and iterate on their own, while you set the scope and review the diffs. Can explain and correct what the agent produced
  • Clear written communication in design docs and PR descriptions

Preferred Qualifications

  • Has shipped an LLM or agent application to real users in production
  • Has run coding agents with more autonomy — letting them carry whole features end to end under constraints you set, or building pipelines where agents build, test and ship with humans designing the system
  • Agent memory: persisting and retrieving user or task context across sessions
  • Feedback loops: turning user feedback and production traces into improved skills, tools or prompts, with human review
  • Context engineering: progressive disclosure, context management for long multi-turn sessions, prompt caching
  • Lightweight decision models (e.g. Jev) for routing, intent classification or caching
  • Guardrails and verification for agents that act on enterprise data
  • Background in predictive modeling, segmentation, or experiment design; understands correlation vs causation
  • Experience with Databricks or Spark; LLM observability tools (e.g. Langfuse)
  • Master’s or PhD degree
  • Industry experience in MarTech and passion for building customer-centric products

#LI-TC1

Similar jobs

Is this your posting and you want it taken down? Email info@careerholo.com.