Site Reliability Engineer – Cloud Infrastructure
Empower Cloud Resilience — Drive Innovation and Reliability at Scale!
Portugal-based opportunity with a fully remote work arrangement (up to 5 days per week).
As a Site Reliability Engineer (AWS), you will own the reliability, scalability, and operational health of our client's AWS-based platform — a leading player in the cloud and infrastructure domain. This is a hands-on engineering role with a systems mindset: you'll define and defend SLOs, drive down toil through automation, lead incident response, and build the tooling and infrastructure-as-code that lets development teams ship safely and fast.
Your main responsibilities:
- Define, measure, and enforce SLOs, SLIs, and error budgets, and use them to drive engineering priorities and release decisions.
- Own the incident lifecycle — detection, response, mitigation, and blameless post-mortems — reducing mean time to recovery through better tooling and runbooks.
- Design and maintain highly available, fault-tolerant AWS architectures across compute, networking, storage, and data layers.
- Build and maintain infrastructure-as-code and CI/CD pipelines that make deployments repeatable, auditable, and low-risk.
- Systematically identify and eliminate toil through automation, replacing manual operational work with self-healing systems.
- Own observability — metrics, logging, tracing, alerting — so problems surface before they reach customers.
- Partner with development teams on capacity planning, performance tuning, cost optimization, and production-readiness reviews.
- Improve the security and compliance posture of the cloud environment in collaboration with security teams.
- Participate in and improve the on-call rotation, tuning alerts to reduce noise and fatigue.
You're ideal for this role if you have:
- Strong production experience operating systems at scale on AWS, including core services (EC2, ECS/EKS, Lambda, VPC, IAM, S3, RDS, CloudWatch).
- Solid grounding in SRE principles — SLOs, error budgets, toil reduction, blameless post-mortems.
- Hands-on expertise with infrastructure-as-code (Terraform preferred; CloudFormation or CDK acceptable).
- Proficiency with container orchestration (Kubernetes / EKS) and CI/CD tooling.
- Strong scripting and automation skills in at least one language (Python, Go, or Bash).
- Deep experience with observability stacks (Prometheus, Grafana, Datadog, ELK, or equivalents).
- Demonstrated ownership of incident response and on-call in a production environment.
- Sound understanding of networking, Linux internals, and distributed-systems failure modes.
Nice to have (not required):
- AWS certifications (Solutions Architect, DevOps Engineer, or SysOps).
- Experience with multi-account / multi-region AWS setups and cost governance.
- Chaos engineering or resilience-testing experience.
- Service mesh, GitOps (ArgoCD / Flux), or policy-as-code (OPA).
- Regulated-industry background (financial services, healthcare).
Language: Fluent English
Eligibility:
- Only candidates with a legal right to work in Portugal or the wider European Union will be considered.
- Only candidates based in Portugal will be considered
#MAKEYourCareerBETTER
Interested? Apply now and include your CV (preferably in English) along with a statement confirming your consent to the processing and storage of your personal data.
Internal number #9753
Benefits
ITDS Clubs
Access to medical insurance
Meal Card
Access to Pluralsight & Udemy
Integrational Events
Seus critérios