Title: Technical Lead-App Development
Area(s) of responsibility
SRE/Middleware & DevOPS engineer
Years of Experience: 10-12 Years
Remote
FTE
We are seeking SRE/Middleware & DevOPS engineer will be responsible for ensuring high availability, relatability and performance of mission -critical production systems. This role combines deep expertise in middleware platforms (Weblogic, Tomcat, Apache and Nginx), CDN technologies (Akami) , Observability tools (New Relic, Splunk) , Automation (Shell Scripting, Terraform) and strong leadership capabilities.
The engineer will lead a team, collaborate with customers, manage production incidents and drive problem management process end-to-end.
Tools: AWS, Nginx, Akamai, Weblogic, Tomcat, Apache, Splunk, New Relic, Shell scripting, Service Now, Bitbucket, Jenkins, Terraform, Docker, Kubernetes, UDeploy and Jira.
Key Responsibilities:
Site Reliability Engineering (SRE)
- Ensure high availability, performance, and resilience of production systems.
- Implement SRE best practices: error budgets, SLIs/SLOs, capacity planning, chaos testing, runbook creation.
- Drive automation to reduce manual operational tasks and improve MTTR.
- Conduct post‑incident reviews (PIRs) and implement long‑term corrective actions.
Middleware & Application Platform Management
- Manage deployments, rollbacks, and environment synchronization across Dev, QA, UAT, and Production.
- Install, configure, upgrade, and maintain WebLogic, Tomcat, Apache, and Nginx servers. Perform JVM tuning, thread pool optimization, connection pool management, and performance tuning.
- Troubleshoot middleware issues related to memory leaks, thread contention, SSL, certificates, and clustering.
Monitoring, Logging & Observability
- Configure dashboards, alerts, and performance insights using New Relic and Splunk.
- Develop log‑based monitoring strategies and anomaly detection.
- Implement proactive monitoring to reduce downtime and improve reliability.
CDN & Edge Platform Management (Akamai)
- Configure Akamai caching rules, WAF policies, edge redirects, and performance optimizations.
- Troubleshoot CDN‑related latency, caching, and routing issues.
- Collaborate with Akamai support for advanced troubleshooting.
Incident, Problem & Change Management
- Lead major incident bridges, coordinate cross functional teams, and provide timely updates.
- Manage problem tickets, root cause analysis, and preventive action plans.
- Ensure compliance with ITIL processes for change, release, and incident management.
Long Description
Leadership & Stakeholder Management
- Lead and mentor a team of SRE/DevOps engineers.
- Provide technical guidance, training, and performance feedback.
- Act as a customer facing technical SME for escalations and production issues.
- Collaborate with product, QA, development, and business teams to ensure smooth delivery.
Behavioral Competencies
- Ownership & Accountability: Takes responsibility for production stability and issue resolution.
- Leadership: Guides team members, manages workload, and drives operational excellence.
- Communication: Clear, structured communication with customers and internal teams.
- Problem Solving: Strong analytical skills and ability to troubleshoot complex issues.
- Collaboration: Works effectively across engineering, QA, product, and business teams.
- Calm Under Pressure: Handles critical incidents with composure and clarity.