Guided by Money Forward AI Vision 2026, Money Forward is driving company-wide AX (AI Transformation) to deliver “digital workers” — AI agents that carry out business operations autonomously. The CDAO Office leads the AI and data strategy that makes this possible across the entire group.
Join the Mepar (Money Forward Engineering Productivity AI Research) team as a Staff AI Engineer and lead the design of production-grade AI agents that power customer-facing products. In this role you’ll own the hardest technical problems at the intersection of agent engineering and backend/infrastructure at scale — turning fast-moving prototypes into reliable, secure, cost-efficient systems that ship to real users.
This is a hands-on, high-leverage individual-contributor role. You’ll set technical direction for how we build, evaluate, deploy, and operate agentic systems — including the durable, long-running execution runtime that lets agents complete real back-office work end to end — and raise the bar for engineering across the team while writing code yourself. You’ll partner closely with product, business, backend, frontend, infrastructure/SRE, and QA teams to deliver impactful, responsible AI end to end.
Responsibilities
- Agentic System Design: Architect, build, and scale multi-agent systems and LLM-powered services for customer-facing products — from prototype to production, built to sustain heavy real-world load.
- Agent Engineering: Design reliable tool-use, function-calling, memory, multi-turn, protocols and modern orchestration frameworks; establish guardrails, evaluation, and safe fallback behavior.
- Backend & Infrastructure at Scale: Own backend services, APIs, and the infrastructure that agents run on — high availability, low latency, secure secrets/credential handling, and horizontal scalability under production traffic.
- Evaluation & Benchmarking (core focus): Own how we measure agent quality. Build the evaluation harness end to end — representative task sets, fixed inputs and reference outcomes, repeatable runners, and scoring across task completion, correctness, latency, cost, and human-intervention rate. Establish LLM-as-a-judge and offline/online eval, wire evaluation into CI as a regression gate, and grow from a pilot task set to a durable benchmark suite that product and QA can trust as an acceptance gate.
- Durable & Long-Running Agent Execution: Design the runtime for agentic tasks that run for ten minutes or longer — durable task lifecycle (submit, status, timeout, cancel, complete, fail), isolated sandboxes for code and file execution, artifact generation and retrieval, and a harness for retries, back-off, checkpointing, resume, and self-recovery.
- Reliability & Cost: Identify bottlenecks; optimize latency, throughput, token/compute cost, and reliability; instrument observability across the agent lifecycle.
- Safety & Governance: Apply security and governance best practices (input validation, content filtering, PII handling, HITL escalation, risk-based sampling) appropriate for customer-facing systems.
- Cross-Functional Technical Leadership & Mentorship: Partner day-to-day with product, business, backend, frontend, infrastructure/SRE, and QA to translate business goals into agent architecture; align standards across BE/AI/Infra, drive system-level decisions that span team boundaries, and mentor engineers on agent development, LLM integration, and evaluation methodology.
Requirements
- 7+ years of professional software engineering experience, with strong recent hands-on delivery (not purely managerial).
- Deep backend engineering expertise — designing, building, and operating large-throughput , low-latency production systems and secure APIs.
- Strong infrastructure skills: cloud ( AWS and/or Azure ), containers ( Docker ), orchestration ( Kubernetes ), Infrastructure as Code ( Terraform ), and CI/CD — including operating and troubleshooting production services.
- Proven experience building AI agents / agentic orchestration for real products, with hands-on use of modern agent frameworks — especially the Claude Agent SDK (agentic loop, tool use, sessions, sandboxed execution, Skills). Comparable depth in LangGraph or similar orchestration frameworks is relevant, as is judgment on when a single agentic loop beats a multi-stage pipeline.
- Strong command of Python or TypeScript for building production services (e.g., FastAPI / FastMCP , async request handling, dependency and lifecycle management, ASGI/Node runtimes)
- Solid understanding of MCP, REST API, GraphQL protocols for tool integration and agent-to-agent communication.
- Demonstrated experience building agent evaluation from scratch — not just consuming dashboards. You have designed task sets and scoring rubrics, run controlled baseline-versus-variant experiments, and used the results to drive architecture decisions.
- Experience with LLM tracing and observability (OpenTelemetry / OpenLLMetry, Langfuse, or equivalent) and with performance and load testing of production services.
- Experience building or operating durable / long-running execution infrastructure — job and task lifecycle management, isolated sandboxes (containers, microVMs, or managed sandbox services), artifact storage, and failure recovery.
- Strong grasp of core CS fundamentals: data structures, algorithms, software design, and engineering best practices.
- English: Business level (equivalent to TOEIC 700 or above)
Nice to haves
While not specifically required, tell us if you have any of the following.
- Experience shipping AI features in customer-facing / B2B SaaS products under real security and compliance constraints.
- Experience with guardrails, safety, and governance frameworks for LLM/agent systems (e.g., OWASP-style threat modeling for agents).
- Experience with cost optimization for LLM usage (context compression, model routing, non-frontier models, gateways).
- Fluent use of AI-assisted development tools (Claude Code, Cursor, GitHub Copilot, Codex) with sound judgment on when to delegate to AI and when to verify.
- Japanese: Not required but nice to have
Compensation
¥11,004,000 ~ ¥20,004,000 annually.