Title: Project Lead-App Development
Area(s) of responsibility
Role Name: Infrastructure Cloud SME
Grade: 5A
Experience: 8 to 10 Years
We are seeking highly experienced Infrastructure Cloud SME to lead the design, engineering, modernization, and operational excellence of large-scale enterprise cloud and platform environments. The ideal candidate will bring deep hands-on expertise across AWS, Pivotal Cloud Foundry (PCF), WebLogic, Akamai, NGINX, Terraform, Splunk, and New Relic, with a strong focus on cloud architecture, infrastructure automation, resilience, performance, security, and production reliability.
This role will serve as a key technical authority for complex infrastructure and cloud initiatives, driving standards, solving critical production challenges, and enabling scalable, highly available, and secure platforms. The candidate should also bring practical knowledge of GitHub Copilot, Model Context Protocol (MCP), and Agentic AI, with the ability to identify and implement AI-assisted opportunities for infrastructure automation, troubleshooting, operational efficiency, and engineering productivity.
Tools: AWS, Nginx, Akamai, Weblogic, Tomcat, Apache, Splunk, New Relic, Shell scripting, Service Now, Bitbucket, Jenkins, Terraform, Docker, Kubernetes, UDeploy and Jira.
Key Responsibilities:
- Serve as the Infrastructure Cloud SME for enterprise cloud and platform environments.
- Design, build, support, upgrade, and troubleshoot AWS and PCF infrastructure.
- Manage AWS services including EC2, VPC, Load Balancers, Route 53, IAM, S3, CloudWatch, Auto Scaling, and related services.
- Support PCF platform infrastructure, application runtime environments, capacity, availability, and platform health.
- Administer and support WebLogic domains, clusters, JVMs, and application deployments.
- Configure and support Akamai and NGINX for CDN, traffic routing, reverse proxy, load balancing, and SSL/TLS.
- Provision and manage infrastructure using Terraform Infrastructure as Code.
- Support cloud networking including DNS, routing, firewalls, security groups, load balancers, certificates, and connectivity.
- Monitor infrastructure and application health using Splunk and New Relic.
- Troubleshoot complex cloud, infrastructure, middleware, network, and performance issues.
- Support high availability, capacity planning, resiliency, disaster recovery, and infrastructure performance.
- Work closely with application, network, security, DevOps, SRE, and cloud teams to resolve infrastructure issues.
- Maintain working knowledge of MCP and Agentic AI for AI-assisted infrastructure operations, troubleshooting, and automation.
Required Skills:
- Experience in Cloud Infrastructure, Infrastructure Engineering, Platform Engineering.
- Strong hands-on experience with AWS and PCF.
- Strong experience with Terraform and Infrastructure as Code.
- Hands-on experience with WebLogic, Akamai, and NGINX.
- Strong experience with Splunk and New Relic.
- Strong knowledge of Linux/Unix, networking, DNS, TCP/IP, load balancing, SSL/TLS, certificates, firewalls, and cloud security.
- Strong troubleshooting, incident management, root-cause analysis, and production support skills.
- Experience supporting large-scale, highly available enterprise cloud environments .
- Knowledge of GitHub Copilot, MCP, and Agentic AI.