Responsibilities
- Design the structure and operational guidelines for managed service offerings across multiple AWS accounts and regions.
- Create foundational infrastructure-as-code components, including Terraform modules and Helm charts, used by other engineers.
- Guide internal observability practices by using the company's own tools to monitor and manage infrastructure performance.
- Lead planning for system capacity, scaling approaches, update rollouts, and cost efficiency in distributed AWS environments.
- Develop platforms and automated systems that increase team effectiveness without requiring linear headcount growth.
- Act as the top-tier technical resource for unprecedented, high-impact customer incidents lacking documented solutions.
- Diagnose and resolve complex system issues involving distributed architectures, Kubernetes, AWS networking, and service mesh technologies under time-sensitive conditions.
- Collaborate directly with customer engineering and operations leaders during extended escalations and architectural redesigns for key accounts.
- Serve as the ultimate incident commander for managed services, providing decisive leadership when advanced expertise is required.
- Develop diagnostic frameworks, runbooks, and tooling that empower mid-level engineers to handle complex situations independently.
- Influence open source strategy within the OpenTelemetry community by initiating and guiding key projects.
- Represent the company in OpenTelemetry special interest groups, shaping the evolution of core components like collectors and exporters.
- Identify gaps in the OpenTelemetry ecosystem that hinder customer success and lead efforts to address them through upstream contributions or internal tools.
- Produce standardized integration guides and reference designs for common deployment environments like Kubernetes and serverless platforms.
- Lead technical contributions to internal open source projects that enhance managed service scalability and reliability.
- Provide final technical validation for complex sales opportunities, including live troubleshooting and architecture review.
- Lead infrastructure and data pipeline discussions for strategic customer accounts in coordination with solutions leadership.
- Conduct in-depth architecture reviews, SLO planning sessions, and instrumentation workshops for complex customer systems, often advising executive stakeholders.
- Lead high-priority proof-of-concept deployments and pilot programs by configuring data pipelines and validating integrations in customer environments.
- Inform product roadmap decisions by identifying critical gaps between current capabilities and the needs of major customers.
- Build internal tools and user interfaces that improve operational efficiency across deployment, monitoring, and rule management.
- Translate field feedback into actionable insights, aligning teams across support, product, engineering, and customer success.
- Balance priorities across multiple domains including managed services, escalation support, open source, and field assistance.
- Coach junior and mid-level engineers on both technical excellence and professional growth, preparing them for advancement.
- Engage with executive stakeholders internally and at customer organizations during critical technical discussions.
Work Arrangement
Remote (Worldwide)