Country/Region:  US
Requisition ID:  38895
Work Model:  Remote
Position Type:  Permanent
Salary Range: 
Location:  US - NEW JERSEY CLIENT SITE

Title:  Technical Lead-App Development

Description: 

Area(s) of responsibility

SRE/Middleware & DevOPS engineer

Years of Experience: 10-12 Years

Remote

FTE

We are seeking SRE/Middleware & DevOPS engineer will be responsible for ensuring high availability, relatability and performance of mission -critical production systems. This role combines deep expertise in middleware platforms (Weblogic, Tomcat, Apache and Nginx), CDN technologies (Akami) , Observability tools (New Relic, Splunk) , Automation (Shell Scripting, Terraform) and strong leadership capabilities.

The engineer will lead a team, collaborate with customers, manage production incidents and drive problem management process end-to-end.

Tools: AWS, Nginx, Akamai, Weblogic, Tomcat, Apache, Splunk, New Relic, Shell scripting, Service Now, Bitbucket, Jenkins, Terraform, Docker, Kubernetes, UDeploy and Jira.

Key Responsibilities:

Site Reliability Engineering (SRE)

  • Ensure high availability, performance, and resilience of production systems.
  • Implement SRE best practices: error budgets, SLIs/SLOs, capacity planning, chaos testing, runbook creation.
  • Drive automation to reduce manual operational tasks and improve MTTR.
  • Conduct post‑incident reviews (PIRs) and implement long‑term corrective actions.

Middleware & Application Platform Management

  • Manage deployments, rollbacks, and environment synchronization across Dev, QA, UAT, and Production.
  • Install, configure, upgrade, and maintain WebLogic, Tomcat, Apache, and Nginx servers. Perform JVM tuning, thread pool optimization, connection pool management, and performance tuning.
  • Troubleshoot middleware issues related to memory leaks, thread contention, SSL, certificates, and clustering.

 Monitoring, Logging & Observability

  • Configure dashboards, alerts, and performance insights using New Relic and Splunk.
  • Develop log‑based monitoring strategies and anomaly detection.
  • Implement proactive monitoring to reduce downtime and improve reliability.

 

CDN & Edge Platform Management (Akamai)

  • Configure Akamai caching rules, WAF policies, edge redirects, and performance optimizations.
  • Troubleshoot CDN‑related latency, caching, and routing issues.
  • Collaborate with Akamai support for advanced troubleshooting.

Incident, Problem & Change Management

  • Lead major incident bridges, coordinate cross functional teams, and provide timely updates.
  • Manage problem tickets, root cause analysis, and preventive action plans.
  • Ensure compliance with ITIL processes for change, release, and incident management.

 

Long Description

Leadership & Stakeholder Management

  • Lead and mentor a team of SRE/DevOps engineers.
  • Provide technical guidance, training, and performance feedback.
  • Act as a customer facing technical SME for escalations and production issues.
  • Collaborate with product, QA, development, and business teams to ensure smooth delivery.

Behavioral Competencies

  • Ownership & Accountability: Takes responsibility for production stability and issue resolution.
  • Leadership: Guides team members, manages workload, and drives operational excellence.
  • Communication: Clear, structured communication with customers and internal teams.
  • Problem Solving: Strong analytical skills and ability to troubleshoot complex issues.
  • Collaboration: Works effectively across engineering, QA, product, and business teams.
  • Calm Under Pressure: Handles critical incidents with composure and clarity.