Job Opportunities for You

Back to All Offers

How to use the job offers’ basket?

Checkout

New

Site Reliability Engineer – Cloud Infrastructure

Portugal
Remote
Angular
English
Expert, Senior
Banking
Infrastructure Specialist

Empower Cloud Resilience — Drive Innovation and Reliability at Scale!

Portugal-based opportunity with a fully remote work arrangement (up to 5 days per week).

As a Site Reliability Engineer (AWS), you will own the reliability, scalability, and operational health of our client's AWS-based platform — a leading player in the cloud and infrastructure domain. This is a hands-on engineering role with a systems mindset: you'll define and defend SLOs, drive down toil through automation, lead incident response, and build the tooling and infrastructure-as-code that lets development teams ship safely and fast.

Your main responsibilities:

  • Define, measure, and enforce SLOs, SLIs, and error budgets, and use them to drive engineering priorities and release decisions.
  • Own the incident lifecycle — detection, response, mitigation, and blameless post-mortems — reducing mean time to recovery through better tooling and runbooks.
  • Design and maintain highly available, fault-tolerant AWS architectures across compute, networking, storage, and data layers.
  • Build and maintain infrastructure-as-code and CI/CD pipelines that make deployments repeatable, auditable, and low-risk.
  • Systematically identify and eliminate toil through automation, replacing manual operational work with self-healing systems.
  • Own observability — metrics, logging, tracing, alerting — so problems surface before they reach customers.
  • Partner with development teams on capacity planning, performance tuning, cost optimization, and production-readiness reviews.
  • Improve the security and compliance posture of the cloud environment in collaboration with security teams.
  • Participate in and improve the on-call rotation, tuning alerts to reduce noise and fatigue.


You're ideal for this role if you have:

  • Strong production experience operating systems at scale on AWS, including core services (EC2, ECS/EKS, Lambda, VPC, IAM, S3, RDS, CloudWatch).
  • Solid grounding in SRE principles — SLOs, error budgets, toil reduction, blameless post-mortems.
  • Hands-on expertise with infrastructure-as-code (Terraform preferred; CloudFormation or CDK acceptable).
  • Proficiency with container orchestration (Kubernetes / EKS) and CI/CD tooling.
  • Strong scripting and automation skills in at least one language (Python, Go, or Bash).
  • Deep experience with observability stacks (Prometheus, Grafana, Datadog, ELK, or equivalents).
  • Demonstrated ownership of incident response and on-call in a production environment.
  • Sound understanding of networking, Linux internals, and distributed-systems failure modes.

Nice to have (not required):

  • AWS certifications (Solutions Architect, DevOps Engineer, or SysOps).
  • Experience with multi-account / multi-region AWS setups and cost governance.
  • Chaos engineering or resilience-testing experience.
  • Service mesh, GitOps (ArgoCD / Flux), or policy-as-code (OPA).
  • Regulated-industry background (financial services, healthcare).


Language: Fluent English 

Eligibility: 

  • Only candidates with a legal right to work in Portugal or the wider European Union will be considered.
  • Only candidates based in Portugal will be considered 


#MAKEYourCareerBETTER

Interested? Apply now and include your CV (preferably in English) along with a statement confirming your consent to the processing and storage of your personal data.
 

Internal number #9753

Benefits

ITDS Clubs

Access to medical insurance

Meal Card

Access to Pluralsight & Udemy

Integrational Events

Your criteria

Want to be notified about new job openings
that match your preferences?

By providing e-mail address and clicking „Notify Me” you agree that ITDS will send e-mail notification that match criteria you have provided.