Site Reliability Engineering (SRE)
Site Reliability Engineer
Purpose of the Role
As a Site Reliability Engineer, you will help improve the stability, reliability, performance, and operational readiness of business-critical digital solutions.
This is not a traditional application support or service management position. You will combine software development, cloud infrastructure, DevOps, integration, and observability practices to identify recurring issues, improve existing solutions, automate operational processes, and prevent incidents before they affect users.
You will work with complex, high-load, and highly integrated front-end and back-end systems in a global technology environment.
Key Responsibilities
Reliability and Software Engineering
- Develop and improve software solutions that increase system stability, reliability, and performance.
- Troubleshoot complex technical issues across applications, integrations, infrastructure, and cloud environments.
- Identify recurring operational problems and implement sustainable technical solutions rather than temporary fixes.
- Refactor and enhance existing code, scripts, services, and automation.
- Contribute to the technical design and continuous improvement of highly integrated digital solutions.
Cloud and DevOps
- Support and improve cloud-based solutions running in AWS environments.
- Work with CI/CD pipelines to automate software build, testing, release, and deployment activities.
- Support containerized applications and orchestration environments using Kubernetes.
- Contribute to infrastructure configuration, secrets management, deployment processes, and operational readiness.
- Work with tools such as Jenkins or comparable CI/CD technologies.
Observability and Monitoring
- Design and implement monitoring, alerting, logging, and observability solutions.
- Proactively identify reliability risks, performance issues, and abnormal system behavior.
- Create meaningful dashboards and alerts that enable teams to detect and resolve issues effectively.
- Work with tools such as Grafana, Elastic Stack, Kibana, Logstash, Dynatrace, Datadog, or comparable observability and application performance management technologies.
- Continuously improve alert quality and reduce unnecessary operational noise.
Integration and API Reliability
- Investigate and resolve issues across integration layers, APIs, gateways, and distributed systems.
- Support the reliability of solutions with multiple internal and external integrations.
- Collaborate with application, platform, infrastructure, and integration teams to diagnose end-to-end system issues.
- Contribute to improvements in API performance, availability, monitoring, and error handling.
- Experience with integration technologies such as TIBCO, Kong, or similar platforms would be beneficial.
Incident Management and Operational Support
- Take ownership of complex incidents and drive them toward resolution.
- Perform root-cause analysis and implement corrective and preventive actions.
- Document technical findings and share knowledge with relevant teams.
- Participate in an on-call rotation after completing the required onboarding and knowledge-transfer period.
- Collaborate with teams across different time zones and participate in a 24/7 on-call rotation, two days per week (within LAM/NAM timezones)
Security and Quality
- Apply security-by-design principles throughout software development and operational activities.
- Support vulnerability identification, remediation, and the implementation of required security controls.
- Ensure that technical changes meet agreed quality, security, and operational standards.
- Contribute to testing, integration validation, deployment readiness, and post-release support.
Key Relationships
- Global IT teams
- Software Engineering and Development teams
- Cloud and Infrastructure teams
- DevOps and Platform Engineering teams
- Integration and API teams
- Information Security
- Business and Product stakeholders
- Globally distributed support and operations teams
What We Are Looking For
- Professional experience in Site Reliability Engineering, Software Engineering, DevOps, Platform Engineering, or a closely related technical field.
- Strong hands-on experience with AWS cloud environments.
- Practical knowledge of CI/CD processes and tools such as Jenkins or comparable technologies.
- Experience with Kubernetes, container orchestration, infrastructure configuration, and secrets management.
- Experience implementing or supporting monitoring, logging, alerting, and observability solutions.
- Practical understanding of application performance management tools such as Dynatrace, Datadog, or similar platforms.
- Strong troubleshooting skills across applications, integrations, infrastructure, and cloud environments.
- Experience working with APIs, gateways, middleware, or complex integration layers.
- Ability to develop or improve code, scripts, automation, and technical solutions.
- Understanding of security, vulnerability management, system testing, release, and deployment practices.
- Strong spoken and written English.
- Ability to collaborate effectively with multicultural and geographically distributed teams.
Nice to Have
- Experience with Elastic Stack, including Elasticsearch, Logstash, and Kibana.
- Experience with Grafana.
- Knowledge of TIBCO, Kong, API gateways, or comparable integration platforms.
- Experience supporting high-load or business-critical digital solutions.
- Previous participation in an on-call rotation or a 24/7 operational model.
- Experience transitioning recurring support activities into automated and scalable solutions.
Professional Skills
- Strong sense of accountability and end-to-end ownership.
- Proactive and self-directed approach to problem-solving.
- Ability to work independently without requiring close supervision.
- Comfortable making technical recommendations and constructively challenging existing approaches.
- Ability to prioritize multiple technical topics in a demanding environment.
- Strong communication skills with both technical and non-technical stakeholders.
- Curiosity and willingness to learn business-specific systems, technical flows, and architectures.
- Long-term interest in developing within the Site Reliability Engineering discipline.
Education and Experience
- University degree in Computer Science, Software Engineering, Information Technology, or a related field, or an equivalent combination of education and professional experience.
- Relevant professional experience in IT, with hands-on exposure to software development, cloud, DevOps, infrastructure, observability, or site reliability engineering.
- Direct SRE experience is highly preferred.
At adidas, we strongly believe that embedding diversity, equity, and inclusion (DEI) into our culture and talent processes gives our employees a sense of belonging and our brand a real competitive advantage.
– Culture Starts With People, It Starts With You –
By recruiting talent and developing our people to reflect the rich diversity of our consumers and communities, we foster a culture of inclusion that engages our employees and authentically connects our brand with our consumers.