Responsibilities
- Support engineering teams across Mimica in developing mature, resilient applications running on GKE, advising on best practices and unblocking issues quickly.
- Develop and maintain infrastructure-as-code using Terraform and ArgoCD, as well as Mimica-specific Platform applications written in Python.
- Manage and maintain GKE environments across multiple regions (EKS experience is also acceptable), including IAM and identity management on GCP (AWS acceptable).
- Investigate and triage issues in existing applications, observe database and system component health, and implement changes to support the SDLC of other teams.
- Participate in a scheduled on-call rotation to monitor system health, respond to incidents, and ensure high platform availability.
- Contribute to migrating our IaC layer from Terraform to Crossplane and support expansion into multi-region clusters and single-tenancy environments.
- Drive FinOps initiatives, helping us get visibility into and control over cloud spend across multiple accounts and environments.
- Help design and implement access patterns for new BYOC and single-tenancy environments, making it seamless for engineers to navigate an increasingly complex multi-environment landscape.
- Instrument, monitor, and improve observability across our stack using tools such as Honeycomb and Grafana.
Requirements
- Solid hands-on experience with Kubernetes, whether on GKE (preferred), EKS, OnPrem, or another managed Kubernetes service. You know your way around cluster configuration, scheduling, RBAC, and debugging live issues.
- Proven experience with Terraform as your primary IaC tool. Familiarity with Terragrunt is a plus.
- IAM and identity management experience on at least one major hyperscaler: GCP (preferred) or AWS.
- Experience with ArgoCD or a comparable GitOps tool for continuous delivery.
- Observability experience with at least one major platform: Honeycomb, Datadog, Grafana, or New Relic. You know how to set up meaningful dashboards, alerts, and traces.
- Comfortable writing code, particularly Python, to automate platform tasks and build internal tooling.
- Background in startups or scale-ups: you are used to ambiguity, moving fast, and wearing multiple hats.
- Strong communication and collaboration skills: You can translate infrastructure complexity into clear language for application engineers, and you know when to raise a flag versus just fixing it.
- Experience leading and participating in incident responses. You systematically decompose problems and know how to get to the root cause without panicking.
Nice to Have
- Experience running GPU workloads on Kubernetes, particularly in the context of ML infrastructure.
- Familiarity with Crossplane (or strong interest in migrating to it from Terraform).
- Experience with Temporal for workflow orchestration.
- Hands-on knowledge of Grafana Alloy, Loki, or similar observability tooling in the Grafana ecosystem.
- Experience with MongoDB, PostgreSQL, or SQL Server.
- Security engineering experience, particularly around hardening Kubernetes environments or supporting compliance initiatives (e.g. SOC 2).
Benefits
- Generous compensation + stock options - aligned with our internal framework, market data, and individual skills.
- Distributed work: Work from anywhere - fully remote, in our hubs, or a mix.
- Company-issued laptop, remote setup stipend, and co-working budget
- Flexible schedules and location
- Ample paid time off, in addition to local public holidays
- Enhanced parental leave
- Health & retirement benefits
- Annual learning & development budget
- Annual workaways and regular virtual & in-person socials
- Opportunity to contribute to groundbreaking projects that shape the future of work
Work Arrangement
Remote (Worldwide)
Additional Information
- Note: Some benefits may vary depending on location and role
- Mimica will only contact candidates from an @mimica.ai email address.
- We do not request banking or sensitive personal information during the recruiting process.