Title: Sr Technical Lead-Cloud & Infra Engg
Area(s) of responsibility
.
Key Responsibilities
1. Ansible Automation & Deployment
- Design, develop, and maintain Ansible playbooks, roles, and collections for provisioning, configuration management, OS patching, and compliance enforcement across enterprise Linux environments.
- Build and manage Infrastructure-as-Code (IaC) pipelines integrating Ansible with CI/CD platforms (Jenkins, GitLab CI/CD) to deliver repeatable, auditable deployments.
- Automate routine Linux operational tasks — user management, package installations, security hardening, disk expansion, log rotation — to reduce manual effort and eliminate configuration drift.
- Develop and maintain Ansible Automation Platform (AAP/AWX) workflows, including dynamic inventories, credentials management, and job scheduling.
- Integrate Ansible automation with AWS APIs and services (EC2, SSM, S3) to support cloud-native Linux provisioning and lifecycle management.
- Version-control all automation artifacts in Git; conduct peer reviews and maintain documentation standards for all playbooks and roles.
2. Linux System Administration & Maintenance
- Administer enterprise Linux distributions (RHEL 7/8/9, CentOS, Oracle Linux, Amazon Linux 2/2023) across on-premises, hybrid, and AWS environments.
- Manage core Linux subsystems: kernel parameters, LVM/storage management, NFS/SMB mounts, network configuration (bonding, VLANs), firewalld/iptables, PAM, SSH hardening, and systemd services.
- Perform security hardening using CIS/STIG benchmarks; manage SELinux policies, sudoers configuration, and PKI/certificate lifecycle.
- Operate and optimize AWS Linux infrastructure: EC2 instance management, EBS volume management, AMI lifecycle, Auto Scaling Groups, and VPC networking for Linux workloads.
- Utilize AWS Systems Manager (SSM) for patch management, parameter store management, session management, and automated runbook execution on Linux fleets.
- Plan and execute scheduled maintenance: OS patching cycles, kernel upgrades, capacity planning, and infrastructure lifecycle activities in compliance with change management processes.
Implement and validate high-availability and disaster recovery configurations including clustering (Pacemaker/Corosync), load balancing, and snapshot-based backup strategies