Staff Software Engineer, Observability

Back to Jobs
Robinhood logo

Staff Software Engineer, Observability

Robinhood

Location

Menlo Park, CA

Experience

Senior

Posted

Jul 30, 2026

Apply by

August 29, 2026

Applicants

0

Early applicantEasy applyFull-timeHybrid

Job Description

## Join us in building the future of finance. Our mission is to democratize finance for all. [An estimated $124 trillion of assets](https://www.cerulli.com/press-releases/cerulli-anticipates-124-trillion-in-wealth-will-transfer-through-2048) will be inherited by younger generations in the next two decades. The largest transfer of wealth in human history. If you’re ready to be at the epicenter of this historic cultural and financial shift, keep reading. ## About the team + role We are building an elite team, applying frontier technologies to the world's biggest financial problems. We're looking for bold thinkers. Sharp problem-solvers. Builders who are wired to make an impact. Robinhood isn't a place for complacency, it's where ambitious people do the best work of their careers. We're a high-performing, fast-moving team with ethics at the center of everything we do. Expectations are high, and so are the rewards. The Observability team's mission is to build and own Robinhood's full-stack observability platform — the foundation that keeps every product, service, and customer experience running reliably at scale. We design and operate the systems that give engineers deep visibility into how Robinhood's infrastructure behaves, ensuring that when something goes wrong, the right people know immediately and can act fast. Our work spans metrics, logs, distributed tracing, and alerting pipelines, and we partner closely with the Robinhood Command Center to ensure our observability systems meet or exceed 99.9% uptime. We believe in monitoring our own monitors — the observability infrastructure is a product, not just a tool. As a Staff Software Engineer on the Observability team, you will be the technical lead shaping the roadmap for how Robinhood observes itself at scale. You will own the observability control plane end-to-end, making architectural decisions that directly impact the reliability and operational health of Robinhood's entire product surface. You'll lead a team of six engineers, drive the strategy for cost-efficient telemetry ingestion, build in-house solutions where off-the-shelf tooling falls short, and integrate AI-driven approaches to accelerate progress toward full-stack observability. This is a high-ownership, high-visibility role where your technical decisions set the direction for the entire organization! This role is based in our Menlo Park, CA office, with in-person attendance expected at least 3 days per week. At Robinhood, we believe in the power of in-person work to accelerate progress, spark innovation, and strengthen community. Our office experience is intentional, energizing, and designed to fully support high-performing teams. ## What you'll do - Define and execute the full-stack observability roadmap, establishing a clear technical vision for metrics, logs, traces, and alerting infrastructure across Robinhood's engineering organization. - Own and evolve the observability control plane, including telemetry ingestion pipelines, cost management strategies, and the tooling that ensures observability components remain highly available. - Lead and collaborate with a team of six engineers to build scalable, self-service observability solutions that enable product and infrastructure teams to move faster with greater confidence. - Establish and maintain SLOs for observability systems that meet or exceed Robinhood's 99.9% uptime target, ensuring the observability platform is as reliable as the services it monitors. - Partner with the Robinhood Command Center and engineering teams across the organization to align on dependency mapping, incident response workflows, and observability standards. ## What you bring - 8+ years of software engineering experience, with a proven track record of owning and delivering large-scale observability or infrastructure platform initiatives. - Deep expertise in Kubernetes and public cloud environments (AWS preferred), with the ability to architect and operate distributed systems at scale regardless of specific vendor tooling. - Strong coding proficiency in one or more languages (Go, Python, or similar) with experience integrating observability agents, libraries, and instrumentation directly into production codebases. - Demonstrated experience owning an observability control plane or telemetry pipeline — including ingestion cost management, cardinality control, and signal routing — in a high-traffic production environment. - Experience with the Vector data pipeline (or equivalent high-throughput log/metrics routing tools) and familiarity with tools such as Prometheus, Grafana, Honeycomb, Humio, or Sentry is a plus. ## What we offer - Challenging, high-impact work to grow your career - Performance driven compensation with multipliers for outsized impact, bonus programs, equity ownership, and 401(k) matching - Top Tier benefits to fuel your work, including 100% paid health insurance for employees with 90% coverage for dependents - Access to the best AI tools on the market and continuous AI skill-building for every employee, technical or not - Lifestyle wallet - a highly flexible benefits spending account for wellness, learning, and more - Employer-paid life & disability insurance, fertility benefits, and mental health benefits - Time off to recharge including company holidays, paid time off, sick time, parental leave, and more! - Exceptional office experience with catered meals, events, and comfortable workspaces. Click [here](https://careers.robinhood.com/benefits) to learn more about our Total Rewards, which vary by region and entity. If our mission energizes you and you’re ready to build the future of finance, we look forward to seeing your application. Robinhood provides equal opportunity for all applicants, offers reasonable accommodations upon request, and complies with applicable equal employment and privacy laws. Inclusion is built into how we hire and work—welcoming different backgrounds, perspectives, and experiences so everyone can do their best. Please review the [Privacy Policy](https://careers.robinhood.com/applicantprivacypolicy) for your country of application.

Key Responsibilities

  • Define and execute the full-stack observability roadmap for metrics, logs, traces, and alerting.
  • Own and evolve the observability control plane, including telemetry ingestion pipelines and cost management.
  • Lead and collaborate with a team of six engineers to build scalable, self-service observability solutions.
  • Establish and maintain SLOs for observability systems to meet or exceed 99.9% uptime targets.
  • Partner with the Command Center and engineering teams to align on dependency mapping and incident response workflows.

Skills Required

KubernetesAWSGoPythonDistributed SystemsObservability Control PlaneTelemetry IngestionSLOsLeadershipProblem solvingCollaborationStrategic thinkingVectorPrometheusGrafanaHoneycombHumioSentry

Benefits

  • Performance driven compensation with multipliers
  • Bonus programs
  • Equity ownership
  • 401(k) matching
  • 100% paid health insurance for employees
  • 90% coverage for dependents
  • Access to AI tools
  • Continuous AI skill-building
  • Lifestyle wallet benefits spending account
  • Employer-paid life & disability insurance
  • Fertility benefits
  • Mental health benefits
  • Paid time off
  • Sick time
  • Parental leave
  • Catered meals
  • Events

App exclusive · Free

Smart Job AI Coach

Your personal interview coach on every job — readiness tips, profile improvements, and role-specific prep. Available only in the Pulse Job app.

Interview readiness

See how prepared you are and what to improve for each role.

Personalized tips

Actionable suggestions based on your profile and the job.

After you apply

Keep coaching momentum from job detail through application success.

Get Smart Job AI Coach in the appFree on iOS and Android