Posted on:August 25, 2026

Senior Site Reliability Engineer at RapidAI

RapidAI is hiring a Senior Site Reliability Engineer in Bengaluru, India. Hybrid.

About RapidAI

RapidAI offers an enterprise clinical AI platform for healthcare providers, with solutions for neurovascular, cardiac and vascular, radiology, and life sciences. The platform is designed to enhance clinical assessment and decision-making and to streamline workflows.

Senior Site Reliability Engineer job description

RapidAI is the trusted leader in deep clinical AI, helping hospitals deliver faster, more informed care through intelligent imaging and integrated workflows. The Rapid Enterprise™ Platform supports disease states across the care spectrum, but it’s our clinical depth that drives the most meaningful impact — improving decision-making, patient outcomes, and health-system performance. Used by more than 2,500 hospitals in over 100 countries and backed by 700+ clinical studies, including research that helped expand national stroke-treatment guidelines, RapidAI is the most clinically validated AI platform in healthcare.

What You Do:

  • Own the availability, performance, and incident response for Rapid's production EKS clusters
  • Design and operate the full observability stack — metrics, logs, traces — with
    Open Telemetry as the foundation
  • Define and track SLOs/SLIs/error budgets; lead post-mortems and drive blameless culture
  • Build and maintain infrastructure-as-code using Terraform, Helm, and GitOps patterns
  • Partner with engineering to bake reliability in early — capacity planning, load testing, chaos engineering
  • Tune autoscaling, networking, and cost efficiency across AWS workloads
  • On-call rotation with the expectation you'll also fix the underlying cause, not just the alert
  •  
    What We Looking For:

  • 10+ years in SRE, DevOps, or infrastructure engineering roles
  • Deep AWS expertise — EKS, EC2, VPC, IAM, RDS, S3, CloudWatch, and the
    surrounding ecosystem
  • Production Kubernetes experience at scale: multi-cluster, multi-tenant, real traffic
  • Hands-on Open Telemetry instrumentation and pipeline ownership (collectors, exporters, backends)
  • Strong foundation in Linux, networking, and distributed systems fundamentals
  • Experience with observability platforms (Prometheus, Grafana, Jaeger, or equivalents)
    Comfortable writing automation in Go, Python, or Bash — you reach for code when the GUI runs out
  • Startup mindset: you make decisions with incomplete information and iterate quickly
  • RapidAI is committed to creating an inclusive and diverse workplace. We provide equal employment opportunities to all employees and applicants and prohibit discrimination and harassment of any type in regard to race, color, religion, age, sex, national origin, disability status, genetics, protected veteran status, sexual orientation, gender identity or expression, or any other characteristic protected by federal, state or local laws.

    Apply now

    Applications go straight to RapidAI. We never sit between you and the employer.

    Apply now ↗
    Get new startup jobs by email
    Pick what you want to hear about. The first email confirms your alert, then we only write when something new matches. You can unsubscribe from any email.
    1 of 15 categories
    How often

    Similar engineering jobs

    All engineering jobs →