All fields marked * are required
Role Summary
The Infrastructure Monitoring Engineer will be responsible for 24x7 monitoring, alert management, incident triaging, and first-level root cause analysis across enterprise IT infrastructure and applications to ensure service availability, performance, and SLA compliance.
Key Responsibilities
路 Perform real-time monitoring of infrastructure and applications using enterprise monitoring tools (APM, NPM, Infra monitoring).
路 Proactively detect, analyze, and respond to alerts related to servers, databases, networks, APIs, and application health.
路 Conduct initial triage and impact assessment of incidents and coordinate with L2/L3 infra, application, and vendor teams.
路 Monitor CPU, memory, disk, network, JVM, database, and API metrics and identify abnormal trends.
路 Validate alerts, reduce false positives, and support alert tuning and threshold optimization.
路 Track incidents end-to-end, ensure timely escalation, and maintain SLA/OLA adherence.
路 Support change, deployment, and maintenance activities from a monitoring readiness perspective.
路 Prepare daily health reports, dashboards, and management summaries.
路 Assist in RCA activities by providing logs, metrics, timelines, and monitoring insights.
路 Support audit and compliance requirements by providing monitoring evidence and reports.
Technical Skills Required
路 Monitoring Tools: Dynatrace, VuNet, AppDynamics, Nagios, Zabbix, or equivalent
路 Infrastructure:
o Servers: Windows Server, Linux
o Network: Basic understanding of TCP/IP, latency, packet loss, VLANs
路 Application Monitoring: JVM, API response times, error rates, service availability
路 Database Monitoring (Basic): Oracle / MS SQL / MySQL (connectivity, performance metrics)
路 ITSM Tools: ServiceNow or equivalent (Incident, Problem, Change)
路 Log Analysis: Kibana / ELK / Splunk (basic)
路 Cloud Exposure (Good to have): Azure / AWS monitoring concept
Required
All fields marked * are required
Role Summary
The Infrastructure Monitoring Engineer will be responsible for 24x7 monitoring, alert management, incident triaging, and first-level root cause analysis across enterprise IT infrastructure and applications to ensure service availability, performance, and SLA compliance.
Key Responsibilities
路 Perform real-time monitoring of infrastructure and applications using enterprise monitoring tools (APM, NPM, Infra monitoring).
路 Proactively detect, analyze, and respond to alerts related to servers, databases, networks, APIs, and application health.
路 Conduct initial triage and impact assessment of incidents and coordinate with L2/L3 infra, application, and vendor teams.
路 Monitor CPU, memory, disk, network, JVM, database, and API metrics and identify abnormal trends.
路 Validate alerts, reduce false positives, and support alert tuning and threshold optimization.
路 Track incidents end-to-end, ensure timely escalation, and maintain SLA/OLA adherence.
路 Support change, deployment, and maintenance activities from a monitoring readiness perspective.
路 Prepare daily health reports, dashboards, and management summaries.
路 Assist in RCA activities by providing logs, metrics, timelines, and monitoring insights.
路 Support audit and compliance requirements by providing monitoring evidence and reports.
Technical Skills Required
路 Monitoring Tools: Dynatrace, VuNet, AppDynamics, Nagios, Zabbix, or equivalent
路 Infrastructure:
o Servers: Windows Server, Linux
o Network: Basic understanding of TCP/IP, latency, packet loss, VLANs
路 Application Monitoring: JVM, API response times, error rates, service availability
路 Database Monitoring (Basic): Oracle / MS SQL / MySQL (connectivity, performance metrics)
路 ITSM Tools: ServiceNow or equivalent (Incident, Problem, Change)
路 Log Analysis: Kibana / ELK / Splunk (basic)
路 Cloud Exposure (Good to have): Azure / AWS monitoring concept
Required