At ZEALS, we’re building the infrastructure behind the next generation of conversational AI products.
Our conversational commerce platform processes millions of customer conversations every day on a high-volume, event-driven system built from dozens of microservices. In parallel, we’re rapidly expanding Omakase.ai, our new AI agent platform, and developing deeper integrations between our products to unlock new capabilities and use cases.
Joining the SRE team means more than maintaining production systems—you’ll help shape the infrastructure that powers multiple AI products. From improving reliability and performance to enabling new feature development and platform integrations, you’ll have the opportunity to work on challenging, large-scale engineering problems while influencing the future architecture of our services.
Responsibilities
- Keep critical customer journeys reliable and available, and respond first when they aren’t.
- Join the on-call rotation for business critical services.
- Own alerting and observability: better coverage, fewer false positives, and catching issues before customers do.
- Build and operate in-house platform tooling: Kubernetes operators, custom Kubernetes APIs, GitOps automation, and deployment tooling that helps teams ship faster and safer.
- Eliminate toil by turning recurring developer requests into self-service, bringing an automation mindset to reduce complexity.
- Drive down platform cost without trading away reliability.
- Design infrastructure that scales with bursty, high-volume traffic, managed as code across many environments.
- Measure the platform against Well-Architected principles and act on what’s weak, whether that’s a reliability gap, an oversized resource, or a manual process worth automating.
Requirements
- 3+ years in software engineering, SRE, infrastructure, or platform engineering on GCP or AWS.
- Hands-on Kubernetes experience: you can debug a workload, not just deploy one.
- Monitoring and observability experience (we use Grafana, Loki, Prometheus, and Thanos).
- Comfortable troubleshooting databases (we run MongoDB, Elasticsearch, and Postgres).
- CI/CD and GitOps experience (we use GitHub Actions, Helm, and Argo CD).
- Able to write and read code for tooling and automation (we use Go, Bash, and Python).
- Eligible to work in Japan, and business-level English.
Nice to haves
While not specifically required, tell us if you have any of the following.
- Terraform and infrastructure-as-code at scale.
- Go proficiency and experience building Kubernetes operators.
- Operating event-driven systems and high-throughput messaging at scale.
- Conversational Japanese.