Senior Observability Engineer
Back to Jobs
MXGet Smart Job AI Coach in the appFree on iOS and Android

Senior Observability Engineer
Location
Chennai, Tamil Nadu, India
Experience
Senior
Posted
Jul 30, 2026
Apply by
August 29, 2026
Applicants
0
Early applicantEasy applyFull-timeWork from Office
Job Description
We are fueled by a moral imperative to advance mankind, and it all begins with our people, our product, and our purpose. Passion isn’t something we turn on and off; it’s woven into everything we do. If you thrive in high-challenge environments, are inspired by exceptional teammates, and are driven to grow beyond what you thought possible, MX is where you belong.
Come build the future with us. Join an award-winning company that isn’t just shaping the financial industry, but transforming it in ways that create meaningful, lasting impact for millions of people.
At MX, reliability is a product. Our infrastructure powers financial applications used by millions of people and processes billions of transactions for major financial institutions, and customers feel every second of downtime.
We're building a new observability function that runs the way we run incident response: the system does the heavy lifting, and people handle judgment, customers, and the exceptions. As a Senior Observability Engineer, you build and operate an observability control plane. You scaffold baselines, score coverage, and turn every real incident into the detection the platform should have caught. This is a multiplier role: you raise the bar for every team through standards and automation instead of building each team's dashboards by hand.
We call it the shepherd model. You shepherd Datadog and partner with our product engineering teams so they observe the right signals for their products. Service owners get real signal instead of noise, and leadership gets coverage and health as a program metric.
This role shares the team pager. Observability and incident response run one on-call roster. You take shifts with the rest of the team and act as Incident Commander when an incident needs one. It is core to the role, not an afterthought.
Engineering at MX runs hybrid infrastructure (AWS and bare metal) with services in Ruby, Go, and Java, messaging over NATS and RabbitMQ, and data on PostgreSQL and Redis. Datadog is our observability platform and incident.io is our incident response platform.
## Job Duties
- Build and operate an observability control plane: automate baseline monitors, dashboards, and tagging standards through the Datadog API and Terraform.
- After significant incidents, produce detection and dashboard gap packs grounded in Datadog and MX investigation patterns, with queries ready to apply.
- Define what "good" looks like for a Ruby, Go, or Java service on Datadog (tags, golden signals, alert quality, dashboard contracts), then audit services against that standard and accept or reject readiness.
- Validate, don't own. Service owners keep their alerts and dashboards; you confirm they are complete and correct, then move on. Escalate to engineering managers when coverage fails or an owner is missing.
- Own the monthly observability and service-catalog health report: departed owners, stale dashboards, services with no monitors, SLO gaps, and coverage trends.
- Run maturity assessments (baseline through SLO, launch-ready, self-serve) and track them over time.
- Tune alerting toward zero false SEV1/2 pages and actionable SEV3/4 alerts, and coach teams on Datadog cost and cardinality.
- Build self-serve onboarding so new services get baseline observability on day one, without a multi-week embed.
- Share the team pager. Rotate on the shared IR & Observability on-call, triage and investigate live incidents with Datadog and MX investigation patterns, and take Incident Commander or supporting technical roles as the incident needs.
- After incidents, close the detection loop (gap packs, new monitors, dashboards) so the pager gets quieter over time.
- Run high-value launch and production-readiness reviews as a checkpoint, not a permanent staffing model.
## Basic Requirements
- BS in Computer Science or equivalent experience
- 5+ years running production observability, SRE, or DevOps. Datadog preferred; strong Grafana/Prometheus, Splunk, or New Relic experience counts if you can ramp on Datadog fast.
- Automation-first engineering in Python, Bash, Go, and/or Terraform, plus Kubernetes proficiency. You encode monitoring standards as code rather than clicking the UI.
- AI- and workflow-literate. You've used or built scripted and AI-assisted workflows to scale reviews, audits, and docs.
- Alerting and SLO strategy: burn-rate and error-budget thinking, with a track record of cutting alert fatigue on evidence.
- Distributed-systems debugging across microservices: latency, connection pools, queues, and cascading failure on Kubernetes and bare metal, with NATS, RabbitMQ, Postgres, and Redis.
- Shared on-call, Incident Commander-capable. You've run or supported incident bridges and written postmortems, and you'll take shifts on the shared IR & Observability rotation.
## Preferred Requirements
- Fintech experience with MX-like architectures
- Google SRE practices: toil elimination, incident management, automation for self-healing
- Cross-functional influence without authority. You've improved teams that don't report to you.
- Governance and reporting: you can produce a monthly health and compliance report leadership reads (orphans, stale entries, gaps, trends).
- OpenTelemetry instrumentation
- Datadog cost optimization at scale (cardinality, log indexing, sampling)
- Incident response platforms (incident.io, PagerDuty, OpsGenie); prior formal Incident Commander experience
- Golang and Ruby on Rails (the MX stack)
## What Success Looks Like
By six months, you're a full participant on the shared on-call rotation and a capable Incident Commander on live SEVs, teams you've engaged have alerts and dashboards that answer "what's broken and where do I look?", and the monthly health report runs largely on its own. By twelve months, incidents get caught earlier because of instrumentation the loop added, new services get baseline observability from a self-serve template on day one, and no team depends on a shepherd for day-one coverage.
Compensation
The expected earnings for this role could be comprised of a base salary and other forms of cash compensation, such as bonus or commissions as applicable.
This pay range is just one component of MX’s total rewards package. MX takes a number of factors into account when determining individual starting pay, including job and level they are hired into, location, skillset, peer compensation.
\\Please note applicants applying for this position must have the legal right to work in India without the need for sponsorship. We are unable to provide work sponsorship for this role, and candidates should be able to verify their eligibility to work in the country independently. Proof of eligibility to work in India will be required as part of the hiring process.
Work Environment
In this role, a significant aspect of the job involves working in the office for a standard 40-hour workweek. We believe that the collaborative nature of our work and the face-to-face interactions among team members are essential for fostering a dynamic and productive work environment. Being present in the office enables seamless communication, facilitates quick decision-making, and encourages spontaneous collaboration that contributes to the overall success of our projects. We value the synergy that comes from having our team members physically together, allowing for immediate problem-solving, idea exchange, and team building.
Key Responsibilities
- Build and operate an observability control plane using Datadog API and Terraform.
- Produce detection and dashboard gap packs after significant incidents.
- Define and audit observability standards for Ruby, Go, and Java services.
- Validate service owner alerts and dashboards for completeness and correctness.
- Own monthly observability and service-catalog health reports.
- Run maturity assessments and track progress over time.
- Tune alerting to reduce false positives and coach teams on cost and cardinality.
- Build self-serve onboarding for new services.
- Share the team pager and act as Incident Commander during live incidents.
- Close the detection loop after incidents to improve future coverage.
- Run high-value launch and production-readiness reviews.
Requirements
- BS in Computer Science or equivalent experience
Skills Required
DatadogPythonBashGoTerraformKubernetesNATSRabbitMQPostgreSQLRedisIncident.ioAWSRubyJavaIncident Commander capabilityShared on-call participationAI and workflow literacyAlerting and SLO strategyDistributed-systems debuggingCross-functional influenceGrafanaPrometheusSplunkNew RelicOpenTelemetryGolangRuby on RailsPagerDutyOpsGenieGovernance and reporting
Benefits
- Bonus or commissions
- Total rewards package
App exclusive · Free
Smart Job AI Coach
Your personal interview coach on every job — readiness tips, profile improvements, and role-specific prep. Available only in the Pulse Job app.
Interview readiness
See how prepared you are and what to improve for each role.
Personalized tips
Actionable suggestions based on your profile and the job.
After you apply
Keep coaching momentum from job detail through application success.


