Job Description
Job Title:  AI Site Reliability Engineer Sr
Posting Start Date:  10/7/26

Standard Job Description

Your Mission:

 

The AI Factory team is seeking an experienced Site Reliability Engineer to lead significant reliability improvements across the platforms and services that support AI development and deployment. You will turn recurring operational problems into scalable engineering solutions through software development, automation, observability, and reliability testing.

 

You will own a defined technical area within the reliability roadmap—from identifying and scoping work through design, implementation, and adoption. Working with platform, application, and operations teams, you will establish reusable approaches that improve reliability across the AI Factory rather than optimizing a single tool or service.

 

Key Responsibilities:
•    Independently lead a significant reliability capability, such as incident-response automation, platform health and observability, release validation, or resilience testing.
•    Identify recurring failure modes and operational toil; design and implement software, automation, and process improvements that address their underlying causes.
•    Build and maintain operational tooling that helps teams validate system health, investigate issues, and recover more effectively.
•    Lead and implement approaches for standardizing and correlating metrics, logs, and traces across services, improving how teams detect and troubleshoot reliability issues.
•    Define and standardize approaches for validating upgrades, releases, and platform changes, including automated testing of failure and recovery scenarios.
•    Define and standardize CI/CD, GitOps, and progressive-delivery practices across teams to reduce deployment risk.
•    Lead technical reviews with partner teams, communicate reliability recommendations, and coordinate implementation across team boundaries.
•    Establish reusable engineering patterns, review code and designs, and mentor less-experienced engineers.
•    Contribute to reliability platforms such as Cluster Concierge while retaining ownership of broader reliability outcomes, not just a single product.

 

Responsible for autonomy hardware and software system integration and testing including verification and validation.Translates customer requirements into product and systems specifications while addressing technical, schedule, and cost considerations; Establishes functional and technical specifications and standards for autonomous systems; Determines sensing hardware components; Recommends and selects appropriate control systems; Integrates and optimizes the autonomous compute and sensing hardware and software; Solves hardware/software interface problems; Develops plan(s) to integrate autonomous functionality into product(s) and platform(s) and other system(s); Develops test plans for validation and verification and procedures for test and evaluation requirements in collaboration with designers and developers; Performs integration testing and coordinates subsystem and/or system testing activities for autonomous programs; Documents and conducts analysis of test results and recommends fixes to software, hardware components, subsystems and systems; Interfaces with other teams involved the development lifecycle for perception

Basic Qualifications

•    Experience designing, developing, and maintaining production software or automation using Go, Python, or a comparable language.
•    Experience operating or engineering Kubernetes-based platforms, including troubleshooting complex service or infrastructure issues.
•    Experience building or improving CI/CD, GitOps, infrastructure automation, or deployment workflows.
•    Experience using observability data—including metrics, logs, or traces—to diagnose problems and improve system health.
•    Demonstrated ability to own a technical initiative through design, implementation, and coordination with other teams.

Desired Skills

•    Experience with OpenShift, GitLab CI/CD, Argo CD, Argo Rollouts, or similar GitOps and progressive-delivery tooling.
•    Experience with Prometheus, Grafana, OpenTelemetry, or comparable observability tools.
•    Experience designing automated reliability tests, upgrade validation, resilience tests, or failure-mode analyses.
•    Experience improving incident response, reducing operational toil, or defining actionable service-health measures.
•    Experience developing internal platforms or tools used by multiple engineering teams.
•    Familiarity with AI/ML platforms, GPU-based infrastructure, or deployments in disconnected environments.
•    Experience mentoring engineers and establishing reusable development or operational practices.
•    Strong oral and written communication skills, and ability to collaborate with cross-functional partners 
•    Creative and resourceful when it comes to problem-solving 
•    Ability to work with internal stakeholders to collect feedback, prioritize tasks, and manage the engineering backlog 
•    Self-motivated, self-directed, and the ability to thrive in a fast-paced environment in an industry that constantly changes

Pay Information

GeoZone Definition: GeoZones are geographic groupings created by Lockheed Martin to align compensation ranges with regional labor markets and cost-of-labor differences across the United States. Locations are assigned a Geo Zone based on the primary work location of the role. 

  • Full-time salary range (GEOZONE 1): $122600.00 - $227600.00 
    • Includes metropolitan areas such as Sunnyvale CA; Pal Alto, CA; New York City metropolitan area; Newark, New Jersey; etc.
  • Full-time salary range (GEOZONE 2): $110300.00 - $204900.00 
    • Includes metropolitan areas such as Denver, CO; King of Prussia, PA; Stratford, CT; Moorestown, NJ; etc.
  • Full-time salary range (GEOZONE 3): $98100.00 - $182100.00
    • Includes metropolitan areas such as Dallas–Fort Worth, TX; Orlando, FL; Grand Prairie, TX; Marietta, GA; etc.
  • Full-time salary range (GEOZONE 4): $88300.00 - $163900.00
    • Includes metropolitan areas such as Camden, AR; Lexington, KY; Lufkin, TX; etc. 

At Lockheed Martin, we know mission success starts with taking care of our people. Our Total Rewards program is designed to attract top talent, support your well-being, and help you grow—both professionally and personally.

The salary range for this position is as listed on the requisition. Please note that the salary information listed is a general guideline only. 
Lockheed Martin considers factors such as (but not limited to) scope and responsibilities of the position, candidate's work experience, education/ training, key skills as well as market(work location) and business considerations when extending an offer.  

Benefits offered: Medical, Dental, Vision, Flexible work arrangements and schedules (e.g., 4x10), 401(k) match, Paid time off, Holidays, Parental Leave, EAP, Flexible Spending Accounts, Education Assistance, Life Insurance, Short-Term Disability, and Long-Term Disability.

  • Annual short-term and/or long-term incentive compensation programs may be offered depending on the position. Payments under these annual programs are not guaranteed and can vary from year to year and are tied to a range of performance metrics.
  • For (Washington state applicants only) Non-represented full-time employees: accrue at least 10 hours per month of Paid Time Off (PTO) to be used for incidental absences and other reasons; receive at least 90 hours for holidays. Represented full time employees accrue 6.67 hours of Vacation per month; accrue up to 52 hours of sick leave annually; receive at least 96 hours for holidays. PTO, Vacation, sick leave, and holiday hours are prorated based on start date during the calendar year.