Overview: Responsible for designing, implementing, and continuously improving highly reliable, scalable, and resilient platform solutions across the enterprise. Operates as a subject matter expert (SME) in Site Reliability Engineering, driving reliability engineering practices, operational excellence, and automation across the Software Development Lifecycle. Leads complex initiatives, influences enterprise engineering standards, and partners with senior stakeholders to improve system stability, observability, and performance. Serves as a mentor and technical leader for less experienced engineers across Technology. Primary
Responsibilities
• Accountable for defining and driving service reliability standards, including SLOs, SLAs, and error budgets across platforms. • Design and implement highly available, fault-tolerant architectures aligned with enterprise scalability and resiliency requirements. • Lead incident management practices, including detection, response, escalation, and recovery processes. • Drive problem management and root cause analysis to prevent systemic issues. • Develop and promote observability strategies, including logging, monitoring, alerting, and tracing. • Lead automation initiatives for self-healing systems and operational workflows. • Contribute to and review technical roadmaps with reliability and performance considerations. • Partner with development, infrastructure, cybersecurity, and architecture teams. • Serve as a technical authority for performance, resilience, and capacity planning. • Drive production readiness practices including performance testing and failover capabilities. • Lead cross-team reliability improvement initiatives. • Participate in and lead post-incident reviews ensuring actionable outcomes. • Mentor engineers on reliability engineering and best practices. • Engage with stakeholders to identify risks and optimization opportunities. • Ensure adherence to risk and regulatory standards and escalate issues when needed. • Maintain internal control standards and compliance expectations. Scope of
Responsibilities
Applies expert-level SRE practices across multiple platforms. Drives enterprise-wide reliability improvements and influences technical direction without direct authority. Supervisory/Managerial
Responsibilities
No supervisory responsibilities.
Education
and
Experience
Required: Associate’s degree and a minimum of 9 years’ systems analysis and/ or application development work experience or Bachelor's degree and a minimum of 7 years' systems analysis and/ or application development work experience. In lieu of a degree, a combined minimum of 11 years’ education and/or relevant work experience, including a minimum of 7 years’ systems analysis and/ or application development work experience. Expert experience in system design, reliability engineering, and production operations. Advanced proficiency in at least one programming or scripting language.
Education
and
Experience
Preferred:
Experience
with observability and incident management tooling.
Experience
with cloud platforms such as AWS or Azure. Strong understanding of CI/CD, DevOps, and SDLC practices.
Experience
defining and implementing SLO/SLI frameworks.
Experience
in regulated environments such as financial services. Strong communication and stakeholder management skills. M&T Bank is committed to fair, competitive, and market-informed pay for our employees. The pay range for this position is $139,700.00 - $232,900.00 Annual (USD). The successful candidate’s particular combination of knowledge, skills, and experience will inform their specific compensation. Location Buffalo, New York, United States of America Great companies have an enduring sense of purpose. At M&T, our purpose is a simple one: make a difference in people’s lives and uplift the communities we serve. M&T Bank Corporation is a financial holding company headquartered in Buffalo, New York. M&T’s affiliates offer advice, guidance, expertise and solutions across the entire financial spectrum, combining M&T Bank’s traditional banking services with the wealth management and institutional capabilities offered by Wilmington Trust. M&T Bank has a network of over 1,000 branches and 2,200 ATMs that span 12 states from Maine to Virginia and Washington, D.C. For more than 165 years, M&T has strived to take an active role in our communities and build long-lasting relationships with our customers. We are a bank for communities—combining the capabilities of a large bank with the care of a locally focused institution. As an employer of choice, we are proud to offer competitive benefits ranging from medical and retirement to forty hours of paid volunteer time, each year. Our core values – integrity, ownership, collaboration, curiosity, and candor – drive the work we do. We seek to further build upon our record of success by bringing in top talent and fresh skill sets while continuing to support the growth and development of all our team members. View M&T’s Human Capital Report to learn more. Ready to join our team? Submit your application today! If you are unable to apply through this site due to technical issues or need an accommodation to apply, please contact us at careersitesupport@mtb.com for assistance. M&T Bank is unwavering when it comes to providing equal employment opportunities to all employees and applicants without regard to race, color, national origin, religion, ethnicity, sex, gender identity, age, disability, citizenship, pregnancy, veteran status, military status, marital status, sexual orientation, genetic information or any other characteristic protected under applicable federal, state or local laws. M&T Bank Corporation has policies and procedures in place to promote a drug free workplace. Career Site Privacy Notice