Skip to main content
Staff AI Engineer (AI Agents Development), CDAO Office, Tokyo
Back to Jobs

Staff AI Engineer (AI Agents Development), CDAO Office, Tokyo

Money ForwardTokyoPosted 0 days ago

Resumen del empleo

Salary
¥11,000,000 - 20,000,000/año
Job Type
Tiempo completo
Japanese Level
No se requiere
Category
Tech & Engineering
Patrocinio de visa

Descripción

**About the company:** Money Forward Minato-ku, Tokyo Money Forward is a fintech startup delivering tools to visualize and improve both individuals'​ and companies'​ financial health. **Responsibilities:** Agentic System Design: Architect, build, and scale multi-agent systems and LLM-powered services for customer-facing products — from prototype to production, built to sustain heavy real-world load. Agent Engineering: Design reliable tool-use, function-calling, memory, multi-turn, protocols and modern orchestration frameworks; establish guardrails, evaluation, and safe fallback behavior. Backend & Infrastructure at Scale: Own backend services, APIs, and the infrastructure that agents run on — high availability, low latency, secure secrets/credential handling, and horizontal scalability under production traffic. Evaluation & Benchmarking (core focus): Own how we measure agent quality. Build the evaluation harness end to end — representative task sets, fixed inputs and reference outcomes, repeatable runners, and scoring across task completion, correctness, latency, cost, and human-intervention rate. Establish LLM-as-a-judge and offline/online eval, wire evaluation into CI as a regression gate, and grow from a pilot task set to a durable benchmark suite that product and QA can trust as an acceptance gate. Durable & Long-Running Agent Execution: Design the runtime for agentic tasks that run for ten minutes or longer — durable task lifecycle (submit, status, timeout, cancel, complete, fail), isolated sandboxes for code and file execution, artifact generation and retrieval, and a harness for retries, back-off, checkpointing, resume, and self-recovery. Reliability & Cost: Identify bottlenecks; optimize latency, throughput, token/compute cost, and reliability; instrument observability across the agent lifecycle. Safety & Governance: Apply security and governance best practices (input validation, content filtering, PII handling, HITL escalation, risk-based sampling) appropriate for customer-facing systems. Cross-Functional Technical Leadership & Mentorship: Partner day-to-day with product, business, backend, frontend, infrastructure/SRE, and QA to translate business goals into agent architecture; align standards across BE/AI/Infra, drive system-level decisions that span team boundaries, and mentor engineers on agent development, LLM integration, and evaluation methodology. Requirements 7+ years of professional software engineering experience, with strong recent hands-on delivery (not purely managerial). Deep backend engineering expertise — designing, building, and operating large-throughput , low-latency production systems and secure APIs. Strong infrastructure skills: cloud ( AWS and/or Azure ), containers ( Docker ), orchestration ( Kubernetes ), Infrastructure as Code ( Terraform ), and CI/CD — including operating and troubleshooting production services. Proven experience building AI agents / agentic orchestration for real products, with hands-on use of modern agent frameworks — especially the Claude Agent SDK (agentic loop, tool use, sessions, sandboxed execution, Skills). Comparable depth in LangGraph or similar orchestration frameworks is relevant, as is judgment on when a single agentic loop beats a multi-stage pipeline. Strong command of Python or TypeScript for building production services (e.g., FastAPI / FastMCP , async request handling, dependency and lifecycle management, ASGI/Node runtimes) Solid understanding of MCP, REST API, GraphQL protocols for tool integration and agent-to-agent communication. Demonstrated experience building agent evaluation from scratch — not just consuming dashboards. You have designed task sets and scoring rubrics, run controlled baseline-versus-variant experiments, and used the results to drive architecture decisions. Experience with LLM tracing and observability (OpenTelemetry / OpenLLMetry, Langfuse, or equivalent) and with performance and load testing of production services. Experience building or operating durable / long-running execution infrastructure — job and task lifecycle management, isolated sandboxes (containers, microVMs, or managed sandbox services), artifact storage, and failure recovery. Strong grasp of core CS fundamentals: data structures, algorithms, software design, and engineering best practices. English: Business level (equivalent to TOEIC 700 or above) Nice to haves While not specifically required, tell us if you have any of the following. Experience shipping AI features in customer-facing / B2B SaaS products under real security and compliance constraints. Experience with guardrails, safety, and governance frameworks for LLM/agent systems (e.g., OWASP-style threat modeling for agents). Experience with cost optimization for LLM usage (context compression, model routing, non-frontier models, gateways). Fluent use of AI-assisted development tools (Claude Code, Cursor, GitHub Copilot, Codex) with sound judgment on when to delegate to AI and when to verify. Japanese: Not required but nice to have Compensation ¥11,004,000 ~ ¥20,004,000 annually. **Requirements:** 7+ years of professional software engineering experience, with strong recent hands-on delivery (not purely managerial). Deep backend engineering expertise — designing, building, and operating large-throughput , low-latency production systems and secure APIs. Strong infrastructure skills: cloud ( AWS and/or Azure ), containers ( Docker ), orchestration ( Kubernetes ), Infrastructure as Code ( Terraform ), and CI/CD — including operating and troubleshooting production services. Proven experience building AI agents / agentic orchestration for real products, with hands-on use of modern agent frameworks — especially the Claude Agent SDK (agentic loop, tool use, sessions, sandboxed execution, Skills). Comparable depth in LangGraph or similar orchestration frameworks is relevant, as is judgment on when a single agentic loop beats a multi-stage pipeline. Strong command of Python or TypeScript for building production services (e.g., FastAPI / FastMCP , async request handling, dependency and lifecycle management, ASGI/Node runtimes) Solid understanding of MCP, REST API, GraphQL protocols for tool integration and agent-to-agent communication. Demonstrated experience building agent evaluation from scratch — not just consuming dashboards. You have designed task sets and scoring rubrics, run controlled baseline-versus-variant experiments, and used the results to drive architecture decisions. Experience with LLM tracing and observability (OpenTelemetry / OpenLLMetry, Langfuse, or equivalent) and with performance and load testing of production services. Experience building or operating durable / long-running execution infrastructure — job and task lifecycle management, isolated sandboxes (containers, microVMs, or managed sandbox services), artifact storage, and failure recovery. Strong grasp of core CS fundamentals: data structures, algorithms, software design, and engineering best practices. English: Business level (equivalent to TOEIC 700 or above) **Nice to have:** While not specifically required, tell us if you have any of the following. Experience shipping AI features in customer-facing / B2B SaaS products under real security and compliance constraints. Experience with guardrails, safety, and governance frameworks for LLM/agent systems (e.g., OWASP-style threat modeling for agents). Experience with cost optimization for LLM usage (context compression, model routing, non-frontier models, gateways). Fluent use of AI-assisted development tools (Claude Code, Cursor, GitHub Copilot, Codex) with sound judgment on when to delegate to AI and when to verify. Japanese: Not required but nice to have **Compensation:** ¥11,004,000 ~ ¥20,004,000 annually.

Requisitos

  • 7+ years of professional software engineering experience, with strong recent hands-on delivery (not purely managerial).
  • Deep backend engineering expertise — designing, building, and operating large-throughput , low-latency production systems and secure APIs.
  • Strong infrastructure skills: cloud ( AWS and/or Azure ), containers ( Docker ), orchestration ( Kubernetes ), Infrastructure as Code ( Terraform ), and CI/CD — including operating and troubleshooting production services.
  • Proven experience building AI agents / agentic orchestration for real products, with hands-on use of modern agent frameworks — especially the Claude Agent SDK (agentic loop, tool use, sessions, sandboxed execution, Skills). Comparable depth in LangGraph or similar orchestration frameworks is relevant, as is judgment on when a single agentic loop beats a multi-stage pipeline.
  • Strong command of Python or TypeScript for building production services (e.g., FastAPI / FastMCP , async request handling, dependency and lifecycle management, ASGI/Node runtimes)
  • Solid understanding of MCP, REST API, GraphQL protocols for tool integration and agent-to-agent communication.
  • Demonstrated experience building agent evaluation from scratch — not just consuming dashboards. You have designed task sets and scoring rubrics, run controlled baseline-versus-variant experiments, and used the results to drive architecture decisions.
  • Experience with LLM tracing and observability (OpenTelemetry / OpenLLMetry, Langfuse, or equivalent) and with performance and load testing of production services.
  • Experience building or operating durable / long-running execution infrastructure — job and task lifecycle management, isolated sandboxes (containers, microVMs, or managed sandbox services), artifact storage, and failure recovery.
  • Strong grasp of core CS fundamentals: data structures, algorithms, software design, and engineering best practices.
  • English: Business level (equivalent to TOEIC 700 or above)

Empleos similares

Explore more in Tokyo

Frequently asked questions

What does Staff AI Engineer (AI Agents Development), CDAO Office, Tokyo at Money Forward pay?
The advertised range is ¥11,000,000–¥20,000,000 per year.
Does this role offer visa sponsorship?
Yes, this position is listed as offering visa sponsorship.
Is this job remote?
This role is listed as on-site. Location: Tokyo. Employment type: full_time.
What level of Japanese is required?
The listing specifies none.