Our client is seeking an Infrastructure Engineer to provide advanced first-line operational support across multiple strategic managed service environments. This role is an exciting opportunity for career progression, as it is not a traditional helpdesk role. The position is full-time and involves participating in a weekly on-call rota to provide 24/7 incident response coverage. The successful candidate will manage monitoring platforms, triage and resolve incidents, validate backup operations, and offer meaningful support for critical incidents across both cloud and on-premises environments. You will work closely within a dedicated five-person team. The role requires strong technical foundations in both Microsoft Azure and VMware virtualization platforms, coupled with a disciplined approach to IT Service Management (ITSM) processes and documentation.
JOB DUTIES:
- Manage and triage alerts from LogicMonitor and Azure Monitor across environments, correlating alerts, closing false positives, and escalating genuine incidents.
- Own the ticket lifecycle for incoming incidents and service requests using ServiceNow, ensuring triage, categorisation, prioritisation, and resolution or escalation. Maintain adherence to service level agreements (SLAs) for responses and updates.
- Perform daily checks on Veeam Backup & Replication job statuses and Azure Backup to log failures, attempt basic remediation, and escalate persistent failures to Senior Engineers.
- Assist Senior Engineers during monthly patch cycles, conducting pre-patch checks, server reboots, and post-patch validations across diverse environments, including Windows and Linux.
- Conduct basic troubleshooting on Azure VMs and VMware vSphere instances, aiming to independently resolve P3/P4 issues.
- Execute basic Windows Server administration tasks, including service restarts, event log analysis, and user access troubleshooting.
- During on-call periods, respond as the first responder for all P1 alerts, performing initial triage, engaging vendor support if necessary, and escalating unresolved issues within stipulated timeframes.
- Maintain and update operational documentation, including runbooks and knowledge base articles, contributing to process improvement initiatives.
JOB REQUIREMENTS:
- A minimum of 3–5 years of experience in enterprise infrastructure support roles, ideally within a managed services environment.
- Proven ability to work with monitoring platforms such as LogicMonitor and Azure Monitor, including experience with alert triage and escalation.
- Hands-on experience with ServiceNow or equivalent ITSM tools for managing the ticket lifecycle.
- Solid grounding in Windows Server administration, including troubleshooting and event log analysis.
- Familiarity with VMware vSphere operations, such as VM console access and basic troubleshooting, as well as Azure VM operations.
- Experience with Veeam Backup & Replication job status checks and Azure Backup monitoring.
- Willingness to participate in a weekly on-call rotation to provide critical incident response.
- Strong communication skills for effective engagement with stakeholders and escalation teams.
- Desirable skills include familiarity with HPE ProLiant, basic Linux fundamentals, and exposure to Azure Arc.
WHAT YOU’LL LOVE:
This role offers a unique opportunity to enhance your technical skills and grow within a supportive team environment. You will be involved in varied responsibilities that allow for professional development and even the potential for progression within the organisation. Your contributions will be essential in maintaining high service standards and ensuring effective incident management while working with advanced technological tools.
Could this be your next move? Submit your CV confidentially to our friendly and dedicated recruitment team by clicking here.