Description
**Request ID:** 257231
**Employee Referral Program – Estimated Payout:** $0\.00
We remain committed to continuing our investment in our employees and supporting your ongoing career development at Scotiabank.
***Purpose***
The Site Reliability Engineer (SRE) ensures the **availability, reliability, scalability, and operational efficiency** of the organization’s critical systems and services by combining software engineering practices with operations.
The SRE works in **close collaboration** with **development, operations, and product teams** to implement and strengthen practices around **observability**, **incident management**, **failure response**, **automation**, and **continuous improvement**, ensuring services meet established **Service Level Agreements (SLA/SLO)** and deliver an **optimal user experience**.
Additionally, the SRE is responsible for **detecting failures in real time**, leading the **initial technical response**, **automating repetitive tasks**, **reducing MTTR**, and providing **data-driven analysis** to prevent future incidents and continuously improve production environment reliability.
#### ***Responsibilities:***
**Service Availability and Reliability**
* Design, implement, and maintain **resilient systems** compliant with **SLO/SLA**.
* Ensure **24/7 operations** and service continuity while respecting **error budgets**.
**Observability and Analysis (end-to-end)**
* Implement and maintain **observability** (metrics, logs, traces) and **actionable alerts**.
* Manage **dashboards** and alert rules in the monitoring platform in use.
* Define, measure, and monitor **SLI/SLO** per service.
* Analyze **trends and degradations** using data (metric, log, and trace queries).
**Incident Management and Postmortems**
* Serve as the **first-level specialized technical responder**: detection and **initial diagnosis**.
**Coordinate escalation** and support **resolution** during **P1/P2 incidents**.
* Document and track **postmortems/RCA** and action plans.
* Reduce **MTTR** and prevent **recurrences**.
**Reliability, Automation, and Continuous Improvement**
* Apply SRE practices (**toil reduction, automation, release readiness, error budgets**).
**Automate** operational tasks (scripts, CI/CD pipelines, remediations).
* Identify and execute **architecture, performance, and cost optimizations**.
**Capacity Management and Scalability**
* Analyze **usage and growth trends** to anticipate infrastructure needs.
* Plan and validate **scalability** and **performance** of services.
**Cross-Functional Collaboration**
* Collaborate with **Development, QA, Security, Infrastructure, and Product** from the design phase.
* Ensure **new services** meet standards for **observability**, **maintainability**, and **reliability** prior to go-live.
**Security and Compliance**
* Ensure compliance with applicable **security, privacy, and regulatory policies**.
* Support **controls, evidence collection, and audits** per internal frameworks.
**Technical Documentation and SRE Culture**
* Maintain clear and up-to-date **documentation** (architecture, processes, runbooks, SLI/SLO, RCA).
* **Promote SRE principles** and best practices across related teams.
#### ***Hierarchical Relationships (job titles only) Primary Manager:***
***(include secondary manager if relevant)***
* Deputy Director, Service Reliability Engineering (SRE)
#### ***Direct Reports:***
n/a
#### ***Shared Reports (solid or dotted line, as applicable):***
* n/a
* Management of high-volume transactional systems operating 24/7\.
* Responsibility for the health and availability of the production ecosystem.
* Generation of executive reports on availability and performance.
* Collaboration with local and global IT teams.
* Improvement of the on-call process.
* Understanding of the Bank’s risk culture and how risk appetite must be considered in daily activities and decisions.
* Ensures compliance with applicable operational and regulatory controls.
* Contributes to reducing operational, regulatory, anti-money laundering, counter-terrorism financing, and conduct risks.
#### ***Education / Experience / Other Information (include only those specific to the role)***
* Bachelor’s degree in Systems Engineering, Computer Science, Telecommunications, or related field.
* Upper-intermediate to advanced English proficiency (spoken and written).
* **5+ years** of experience in production environments requiring **high availability** and high transaction volume (24/7 operations\).
* **3+ years** supporting production or roles related to reliability, operations, or monitoring.
* **4+ years** of experience in cloud engineering (**AWS, GCP, Azure**) or equivalent functions.
* Experience designing, implementing, and maintaining **SLI/SLO** and SRE practices.
* Experience with **microservices**, container-based workloads, and serverless functions.
* Experience in designing **resilient, scalable, and secure architectures**.
* Participation in **complex incident management**, detailed diagnostics, and root cause analysis.
* Proven ability to proactively identify issues, bottlenecks, and improvement opportunities.
At Scotiabank, we value the unique skills and experiences each person brings to the Bank and are committed to creating and maintaining an inclusive and accessible environment for everyone. All employees must comply with the Bank’s policies, standards, codes, and guidelines related to non-discrimination and workplace accommodations. If you require any accessibility accommodation during the hiring process, please inform our Talent Attraction team\*\*Scotiabank is an inclusive employer that respects diversity and does not discriminate in any way\*\*\*\*Under no circumstances does Scotiabank request pregnancy or HIV testing\*\*Thank you for your interest. However, only candidates selected for interviews will be contacted.
Location(s): Mexico : Mexico City : Cuauhtémoc
Scotiabank is a leading bank in the Americas. Inspired by our corporate purpose — “for tomorrow” — we help our clients, their families, and their communities succeed through a full range of advice, products, and services in personal and commercial banking, wealth management, private banking, corporate and investment banking, and capital markets.
At Scotiabank, we value the unique skills and experiences each person brings to the Bank and are committed to creating and maintaining an inclusive and accessible environment for everyone. If you require an accommodation (e.g., an accessible interview location, documents in alternate formats, sign language interpretation, or assistive technology), please let our Recruitment team know during the recruitment and selection process. For technical support, click here. Candidates must apply directly online if they wish to be considered for this position. We thank all candidates for their interest in this career opportunity at Scotiabank; however, only those selected for an interview will be contacted.