Elsevier logo

Elsevier

Senior Site Reliability Engineer I

🇺🇸 Philadelphia, Pennsylvania 🕑 Full-Time 💰 $95K - $159K 💻 Information Technology 🗓️ September 30th, 2026
Terraform AWS CI/CD

Edtech.com's Summary

Elsevier is hiring a Senior Site Reliability Engineer I to keep critical platforms and services reliable, scalable, and performant as part of an Embedded Innovation Team. The role leads complex reliability initiatives and automation work that cuts operational toil, drawing on observability, incident response, and distributed systems expertise. The engineer also supports handover so the solutions stay owned and operable after the squad moves on.

Highlights
  • Create monitoring queries and establish service level baselines.
  • Support senior engineers during incidents and contribute to post-mortems and RCAs.
  • Participate in disaster recovery tests and test availability, reliability, and recoverability in non-production environments.
  • Implement automation and execute code in production environments.
  • Support the deployment, monitoring, and reliability of services that integrate AI tools.
  • Contribute to SRE knowledge documentation and infrastructure topology drawings.
  • Bring advanced Terraform expertise and hands-on AWS operations across ECS, RDS, ALB, VPC, IAM, and Lambda.
  • Build and troubleshoot GitHub Actions CI/CD workflows and work with Docker and ECS Fargate containers.
  • Use Linux, Git, and Bash/Python scripting for automation, alongside incident response and observability skills.
  • Base pay range of $95,300 - $158,800 plus eligibility for an annual incentive bonus.

Senior Site Reliability Engineer I Full Description

Locations:
Philadelphia, PA (Market St); Home based-Massachusetts
Time type:
Full time
Job requisition id:
R119025

Senior Site Reliability Engineer
Are you passionate about building resilient, scalable systems that power mission-critical applications?
Do you thrive on automating operations, improving reliability, and ensuring exceptional system performance?

About the team:
Embedded Innovation Teams are cross-functional squads embedded within our segments to rapidly turn internal AI experimentation into validated, reusable solutions, building the capabilities we need to deliver customer value and growth. We work problem-first rather than tool-first, directly inside segment and function teams, improving the internal workflows that help our people deliver better outcomes for customers, faster.

About the role:
As a Senior Site Reliability Engineer (SRE), you will play a key role in ensuring the reliability, scalability, and performance of our critical platforms and services. You will lead complex reliability initiatives, drive automation efforts to reduce operational toil, and help build resilient systems that deliver exceptional customer experiences.
You will leverage your expertise in observability, incident response, and distributed systems to proactively identify and resolve reliability challenges. Working closely with engineering teams, you will design and implement solutions that improve service availability, streamline operations, and enhance system recovery capabilities.
You will hold a high bar on code quality, flag risks and blockers early, and work alongside host-function stakeholders to make sure what you build fits real workflows, not assumed ones. You will also support handover and capability-building so the solution is owned and operable after the squad moves on.

Key Responsibilities:
  • Creating monitoring queries and establishes service level baselines.
  • Supporting senior engineers during incidents.
  • Making contributions during post-mortems and RCAs.
  • Participating in disaster recovery tests.
  • Implementing automation and executes code in production environments.
  • Contributing to SRE knowledge documentation.
  • Supporting the deployment, monitoring, and reliability of services integrating AI tools.
  • Supporting architecture and senior engineers in the creation of infrastructure topology drawings and deployment workflows.
  • Carrying out the testing of availability, reliability, and recoverability in non-production environments.

Requirements:
  • Advanced Terraform: Expertise in modules, providers, state management, lifecycle controls, drift detection, safe refactoring, and remote state (S3, locking, cross-stack dependencies).
  • AWS Operations: Hands-on experience managing production, multi-account, multi-region AWS environments across ECS, RDS, ALB, VPC, IAM, Route53, ECR, S3, Lambda, DynamoDB, SQS, Secrets Manager, KMS, and CloudWatch.
  • GitHub Actions CI/CD: Experience building and troubleshooting reusable workflows, OIDC authentication, approval gates, runners, Terraform deployments, application deployments, and migration pipelines.
  • ECS Fargate & Containers: Knowledge of Docker, ECR, ECS task definitions/services, IAM roles, health checks, autoscaling, ALB integration, and deployment rollbacks.
  • AWS Networking & Security: Proficiency in VPCs, networking, ALBs, Route53, ACM/TLS, IAM, OIDC, Secrets Manager, KMS, and cloud security best practices.
  • Incident Response & Observability: Skilled in troubleshooting using logs, metrics, alarms, deployment history, root cause analysis, rollback decisions, and operational runbooks.
  • Linux & Automation: Strong Linux and Git fundamentals with Bash/Python scripting for AWS CLI automation, CI/CD, and operational tooling.
  • AI Tooling Deployment: Hands-on experience integrating and operating AI services and APIs in production, including monitoring, reliability, and security practices for AI-powered features.
  • Developer Enablement: Ability to support multiple engineering teams, troubleshoot across infrastructure and application layers, document solutions, and enable secure self-service practices.

U.S. National Base Pay Range: $95,300 - $158,800. Geographic differentials may apply in some locations to better reflect local market rates.
This job is eligible for an annual incentive bonus.

We know your well-being and happiness are key to a long and successful career. We are delighted to offer country specific benefits. Click here to access benefits specific to your location.
We are committed to providing a fair and accessible hiring process. If you have a disability or other need that requires accommodation or adjustment, please let us know by completing our Applicant Request Support Form or please contact 1-855-833-5120.
Criminals may pose as recruiters asking for money or personal information. We never request money or banking details from job applicants. Learn more about spotting and avoiding scams here.
Please read our Candidate Privacy Policy.
We are an equal opportunity employer: qualified applicants are considered for and treated during employment without regard to race, color, creed, religion, sex, national origin, citizenship status, disability status, protected veteran status, age, marital status, sexual orientation, gender identity, genetic information, or any other characteristic protected by law.
USA Job Seekers:
EEO Know Your Rights.