Skip to content
McCoy
JPMorgan Chase

Site Reliability Engineer III

JPMorgan Chase·Engineering·IN
Posted Sep 11, 2026·Open for 10 days (and counting)
Hiring for this role?

Generate a McCoy IQ challenge in 30 seconds.

See how candidates think and approach the work this role demands, before the phone screen. We'll build a video challenge from this posting, and you can edit or share it before it goes live.

Key details

Function
Engineering
Seniority
Mid
Workplace
On-site
Location
IN
Specialty
Devops Sre
Tech stack
GrafanaDynatracePrometheusDatadogSplunkPythonPyspark

Job Description

Join a dynamic team where your expertise in site reliability engineering will shape the future of AI/ML data platforms. Unlock opportunities for growth and impact as you help build resilient, market-leading solutions.


As a Site Reliability Engineer III at JPMorgan Chase within the AI/ML Data Platforms team, you will play a pivotal role in developing scalable and resilient data solutions. You will engage in root cause analysis, production changes, and strategic initiatives that drive operational excellence. Your experience will help mentor team members and foster collaboration across global teams. Together, we create innovative solutions that support the firm’s mission and community.

 

Job responsibilities

  • Build and support scalable, resilient AI/ML data solutions
  • Coordinate incident management coverage for effective application issue resolution
  • Collaborate with cross-functional teams to perform root cause analysis and implement production changes
  • Develop and support AI/ML solutions for troubleshooting and incident resolution
  • Mentor and guide team members to drive strategic change
  • Manage budgetary considerations and staffing challenges
  • Uses enterprise-authorized AI capabilities within the work environment to accelerate incident triage, troubleshooting, and post-incident analysis, validating outputs and handling operational data according to sensitivity and security requirements.
  • Applies enterprise-authorized AI capabilities within the work environment to identify patterns in operational signals that indicate reliability risk or recurring toil, prioritizing reuse-first improvements tied to SLO outcomes.
  • Partner with colleagues across global teams to deliver impactful results

 

Required qualifications, capabilities and skills

  • Formal training or certification on site reliability engineering concepts and 3+ years applied experience 
  • Proficient in site reliability culture and principles, with familiarity in implementing site reliability within an application or platform
  • Proficiency in running production incident calls and managing incident resolution
  • Experience in observability including white and black box monitoring, service level objective alerting, and telemetry collection using tools such as Grafana, Dynatrace, Prometheus, Datadog, Splunk, and others
  • Working knowledge of using enterprise-authorized AI capabilities within the work environment to support SRE workflows with strong validation habits and awareness of data sensitivity
  • Ability to validate AI-assisted operational recommendations before applying changes, escalating when uncertain and following data sensitivity requirements
  • Strong understanding of SLI/SLO/SLA, Error Budgets and Proficiency in Python or PySpark for AI/ML modeling
  • Must be able to reduce toil by building new tools to automate repeated tasks
  • Hands-on experience in system design, resiliency, testing, operational stability, and disaster recovery
  • Awareness of risk controls and compliance with departmental and company-wide standards
  • Ability to work collaboratively in teams and build meaningful relationships to achieve common goals

 

Preferred qualifications, capabilities and skills

  • 4+ years in an SRE or production support role with AWS Cloud, Databricks, Snowflake or similar technologies
  • AWS and Databricks certifications
 
Audit details(provenance, verification trail, raw fields)

Core fields

Posting ID
jpmorgan:210789842
Title
Site Reliability Engineer III
Function
Engineering
Location
Hyderabad, Telangana, India
Workplace mode
unspecified
Posted at
2026-09-11 05:42:52Z
Compensation
undisclosed

Provenance

First seen (our tracker)
2026-09-11 09:52:56Z
Last seen
2026-09-13 09:44:21Z
Last updated
2026-09-13 09:44:21Z
Removed at
still open
Days open
Open for 10 days (and counting)
ATS adapter
oracle_hcm
ATS slug
jpmc.fa.oraclecloud.com|CX_1001

Verification trail

  1. still_live2026-09-21 10:21:22Z
    via oracle_hcm
    evidence
    {
      "url": "https://jpmc.fa.oraclecloud.com/hcmUI/CandidateExperience/en/sites/CX_1001/job/210789842"
    }

LLM enrichment

Enriched at 2026-09-11 20:23:31Z. Enrichment runs once per posting, never re-runs.
Seniority
ic_l3
Role archetype
engineering
Specialty
devops_sre
Workplace mode
unknown
City (normalized)
Country (normalized)
United States
Comp range
Tech stack
grafanadynatraceprometheusdatadogsplunkpythonpyspark
Novel role archetype?
no

See how we measure for definitions, or our corrections log for known issues. Found something wrong? Flag a correction.