Title: Project Lead-App Development
Area(s) of responsibility
Role Summary
As Observability Engineer / SME, you will be the technical authority for designing, implementing, and operating enterprise monitoring and observability solutions across hybrid IT environments. Working hands-on with platforms such as HPE OpsRamp and SolarWinds (and similar ITOM/observability tools), you will deliver end-to-end monitoring — onboarding, discovery, service mapping, alert correlation, and ITSM-integrated event management — to improve reliability, reduce noise, and accelerate MTTD/MTTR for our clients. You will also mentor L1/L2 teams and serve as the central point of expertise for monitoring delivery.
Key Responsibilities
Design & Implementation
- Design, deploy, configure, customize, and integrate observability solutions using HPE OpsRamp, SolarWinds, and similar platforms (Nagios, Zabbix, ManageEngine OpManager, etc.) across on-prem, cloud, and edge infrastructure.
- Analyze client infrastructure design, architecture, and security standards to define fit-for-purpose monitoring solutions.
- Develop and maintain architecture blueprints, design documents, SOPs, and operational runbooks for monitoring and event-management systems.
- Lead onboarding, discovery, service mapping, alert correlation, and runbook automation for green-field and brown-field engagements.
Operations & Event Management
- Own day-to-day operations of observability platforms — ensuring high availability, performance, capacity, data accuracy, and timely upgrades/patching.
- Manage event-management processes: event ingestion, correlation, noise reduction, deduplication, and routing to the appropriate resolver groups.
- Troubleshoot complex incidents escalated from L1/L2 teams; lead root-cause analysis (RCA) and turn telemetry into actionable reliability insights.
- Establish and continuously improve standards, thresholds, and SLA-based response models; maintain telemetry data hygiene (retention, indexing, access controls).
- Integrate monitoring tools with ITSM platforms (ServiceNow, BMC, Ivanti) and cloud platforms (AWS, Azure, GCP) — including alert/ticketing workflows via MID server/connectors.
Required Skills & Experience
Monitoring & Observability (mandatory)
- 6–8 years of hands-on experience in IT monitoring tool deployment, implementation, operational support, and report building.
- Strong hands-on expertise with HPE OpsRamp and SolarWinds and their modules; experience with similar tools (ManageEngine OpManager, Nagios, Zabbix, PRTG) is an advantage.
- Working knowledge of broader observability/APM platforms — Dynatrace, AppDynamics, Datadog, Prometheus, Grafana, ELK, Splunk — and OpenTelemetry concepts.
- Proven experience in Managed Services / NOC delivery and ITSM-integrated event management.
Infrastructure, Cloud & Automation
- Strong understanding of servers, networks, virtualization, storage, and hybrid cloud environments.
- Hands-on integration experience with ITSM tools (ServiceNow preferred, BMC, Ivanti) and cloud platforms (AWS, Azure, GCP).
- Scripting for automation and integration — Python, Bash, or PowerShell; familiarity with containerization (Docker, Kubernetes) is a plus.
Sound understanding of ITIL processes — Incident, Problem, Change, and Configuration Management