I&IT ONLY: Metrolinx's Innovation and Information Technology group supports female team members via "Go Tech Women" an affinity group for women in Information Technology, led by our Chief Information Officer.
If you enjoy technology and innovation, value diversity, appreciate work/balance and are looking for an opportunity to make a better world via public service, Metrolinx would like to hear from you!
The SRE and Operations embed resilience, recoverability, and operational excellence across Metrolinx I&IT services. Through continuous delivery, infrastructure as code, proactive observability, and tested disaster recovery capabilities, the function reduces operational risk and enables rapid recovery from major outages, cyber incidents, and disruptions.
What will I be doing?
- Drive SRE & Operations Leadership: Lead, coach, and direct engineering teams responsible for day-to-day infrastructure operations, platform reliability, data center management, and disaster recovery.
- Hands-On Leadership Role: Operate as a hands-on technical leader who remains actively engaged in architecture, engineering decisions, operational problem-solving, incident response, and delivery execution while leading and developing the team.
- Establish Operating Models & Standards: Build and govern the organization’s SRE operating model, priorities, and cloud reliability standards covering availability, scalability, recoverability, and operational support across public and hybrid cloud environments.
- Govern Cloud & Infrastructure Reliability: Define production readiness criteria and operational acceptance guardrails for new or materially changed services, ensuring architecture, monitoring, logging, alerting, security controls, and recovery automation are fully validated prior to release.
- Manage Service Health: Define Service-Level Indicators and Service-Level Objectives for critical services to balance deployment velocity with platform stability.
- Direct Incident Response & Post-Incident Learning: Lead major incident response efforts to restore disrupted services rapidly and facilitate post-incident reviews to ensure root causes are remediated and tracked through engineering backlogs.
- Day-to-Day Operations: Oversee day-to-day data centre operational management, including HPE Synergy, Hyper-V, VMware, SAN/NAS storage configuration, F5, NetScaler, file servers, database enablement (Oracle Data Guard, SQL Always On, Oracle ExaCC), data backup, networking integration (Cisco ACI and Firepower, Palo Alto Firewalls) and physical infrastructure tuning.
- Proactively Manage Infrastructure Risk: Direct reliability risk assessments, capacity forecasting, and controlled resilience testing including but not limited to multi-region, availability-zone, network, and identity failure scenarios to eliminate single points of failure.
- Steward Budgets: Support the governance and allocation of an annual capital budget, managing operational and capital budgets to meet product roadmap objectives within financial constraints.
- Govern Vendor Relationships: Direct vendor deliverables, evaluate technology proposals, participate in contract negotiations, and establish Service Level Agreements (SLAs) and Operational Level Agreements (OLAs) with strategic partners.
- Build a Culture of Innovation & Excellence: Build high-performance team capabilities in automated operations, infrastructure as code, and SRE disciplines.
Flexible Hours: Preparedness to occasionally work outside regular business hours including evenings and weekends during emergency major incidents or urgent technology deployments.
Education Completion of a degree in Computer Science, Information Systems, Business Administration, Engineering, or a related discipline.
Certifications
- Microsoft Certified Azure Solutions Expert, Azure DevOps Engineer, or Equivalent is an asset.
- Cisco Certified Network Professional (CCNP) is an asset.
- Certified Information System Security Professional is an asset.
- Certified Kubernetes Administrator is an asset.
- Information Technology Infrastructure Library (ITIL) Foundations is an asset.
- Agile certification (Agile Certified Professional, Certified Scrum Product Owner) is an asset.
- Project Management Professional (PMP) or PRINCE2 is an asset.
Professional Experience & Technical Mastery
- Progressive Leadership: Demonstrated years of progressively responsible experience in infrastructure operations, platform engineering, cloud engineering, or SRE within a large-scale service delivery environment, including leadership of technical teams and complex operational programs.
- Large-Scale Azure Cloud Migration: Hands-on experience leading large-scale cloud migrations from on-premises environments to Microsoft Azure, including migration planning, architecture, execution, risk management, cutover, and operational transition.
- SRE & Cloud Practice: Deep knowledge of Site Reliability Engineering frameworks, dynamic resource management, service quotas, autoscaling, infrastructure as code (IaC), policy as code, and cloud cost optimization.
- Relationship Management: Exceptional communication, facilitation, and negotiation skills to brief executive leadership, influence internal stakeholders, manage vendor contracts, and introduce modern methodologies (DevOps, Lean, Agile).
Core Technologies (Including but not limited to)
- Cloud Platforms: Microsoft Azure, Copilot, Exchange Online, Teams, SharePoint, etc.
- DC Infrastructure: VMware, Hyper-V, HPE compute and storage, and Brocade SAN.
- Network & Application Delivery: Palo Alto Firewalls, F5, Citrix NetScaler, Fortinet SD-WAN, and Infoblox.
- Identity & Security: Microsoft Active Directory, Zscaler, Microsoft Entra ID, Okta, Microsoft Defender, and CyberArk Privilege Cloud.
- Cloud Operations, Automation & Observability: Grafana, ServiceNow, Terraform, Ansible, and GitHub.
- Operating Systems & Management Platforms: Red Hat Enterprise Linux, Windows Server, SCVMM, SCCM, and Red Hat Satellite.
- DR Solutions: Hands-on experience implementing enterprise recovery technologies, including Azure Cloud DR, Cisco ACI multisite, Zerto, Hyper-V, Oracle Data Guard, SQL Always On, F5 application delivery gateway, and related replication solutions.
Process Acumen: Understanding of project management practices, ITIL frameworks, risk and business impact analysis (BIA), etc.
Metrolinx is an equal opportunity employer and committed to a diverse and inclusive workforce. We are also committed to offering reasonable accommodation to job applicants in accordance with applicable legislation. If you require assistance or an accommodation at any point during the hiring process, please contact us at: 416-202-5601 or email [email protected].