The Senior Observability Engineer is responsible for defining, leading, and advancing enterprise observability strategy, architecture, and implementation across applications, platforms, infrastructure, and operational services. This role requires deep hands-on expertise with Dynatrace SaaS and recent Dynatrace platform capabilities and innovations, as well as strong experience with OpenTelemetry, cloud-native observability, AIOps, and agentic operations.
This individual will serve as a principal-level technical leader, applying strategic thinking, systems architecture expertise, and sound decision-making to establish modern observability standards and scalable telemetry practices. Working in a fast-paced, highly collaborative technology environment, this role partners closely with engineering, site reliability, platform, security, infrastructure, operations, and architecture teams to guide the organization beyond traditional monitoring toward more intelligent and adaptive operational models.
Essential Duties & Responsibilities:
-
Define and evolve our enterprise observability vision, standards, principles, and roadmap, using strategic thinking and sound judgment to align technical direction with business needs.
-
Design observability solutions for Kubernetes, containers, microservices, distributed applications, and public cloud environments.
-
Apply AIOps and agentic operations capabilities to strengthen detection, event correlation, diagnosis, automation, operational response, and continuous improvement.
-
Develop and maintain service health models, SLOs, dashboards, alerting strategies, telemetry governance, and observability best practices that support operational excellence and service reliability.
-
Collaborate across engineering, site reliability, platform, security, infrastructure, operations, and architecture functions to expand observability adoption and maturity.
-
Guide the transition from legacy monitoring practices to modern, adaptive, and outcome-focused observability models, including initiatives involving production systems and enterprise platform transformation.
-
Mentor engineers, influence technical direction, solve complex problems, and help shape enterprise architecture and engineering standards.
Required Qualifications :
-
Master’s degree from an accredited college or university in Computer Science, Information Systems, Engineering, or a related technical field.
-
8+ years of experience in observability, monitoring, site reliability engineering, platform engineering, infrastructure engineering, or related technical disciplines.
-
Strong understanding of metrics, logs, traces, event correlation, service health, alerting models, and operational intelligence.
Preferred Qualifications:
-
Strong knowledge of site reliability engineering principles, including SLIs, SLOs, incident response, and operational resilience.