AWS Cloud Operations Engineer
Description du Poste
AWS Cloud Operations Engineer
Managed AWS Infrastructure and Platform Support
|
Job family |
Cloud Engineering / Cloud Operations |
|
Role type |
Hands-on L2/L3 managed cloud operations |
|
Location |
India / Remote or hybrid, subject to demand approval |
|
Coverage |
24x7 hours with rotational on-call and planned maintenance |
|
Experience guide |
5–10 years overall, including 3+ years in AWS operations |
|
Reporting |
Cloud Operations Lead / Technical Account Manager |
Position Overview
Provide hands-on operations support for secure, reliable and cost-aware AWS environments. The engineer manages production and non-production resources, monitoring, incidents, security controls, maintenance, governance, documentation and continual improvement while working within approved architecture and operational standards.
Key Responsibilities
· Build and support on-demand AWS infrastructure within approved patterns, including AWS accounts and Organizations, VPC, EC2, EBS, S3, IAM, Route 53, Elastic Load Balancing, Auto Scaling, CloudWatch, CloudTrail, AWS Config, Systems Manager, Backup and KMS.
· Maintain inventory, ownership, configuration visibility, tagging and platform health across managed environments.
· Continuously monitor availability, performance, capacity, service health, logs and operational risk; respond to alerts and restore service.
· Own incident triage and resolution, contribute to major incident management, complete RCA and track preventive actions.
· Operate cloud identity, access, network and security controls in accordance with least privilege, policy and audit requirements.
· Support OS patching, platform updates, scheduled reboots, backup checks, restore tests, certificate/secret lifecycle and infrastructure maintenance.
· Detect configuration drift, policy violations, vulnerabilities and unsupported resources; coordinate remediation through controlled change.
· Use automation and Infrastructure as Code to reduce manual work, improve consistency and create repeatable operational outcomes.
· Support disaster recovery readiness, continuity testing, recovery documentation and operational handover.
· Participate in service reviews, stakeholder escalations, audit requests and continuous service improvement initiatives.
· Identify utilization and cost optimization opportunities without compromising resilience, performance or security.
Experience and Qualifications
· Bachelor’s degree or equivalent practical experience in IT, Computer Science, Engineering or a related field.
· Hands-on experience administering AWS infrastructure in enterprise production environments.
· Working knowledge of Windows and/or Linux operating systems, TCP/IP, DNS, routing, firewalls, certificates and backup concepts.
· Experience with ITSM processes, controlled change, incident management and operational documentation.
· Ability to join rotational on-call support and planned maintenance windows.
Preferred Certifications
· AWS Certified CloudOps Engineer - Associate
· AWS Certified Solutions Architect - Associate (preferred)
· AWS Certified Security - Specialty (preferred for senior roles)
· ITIL Foundation
Core Competencies
· Strong troubleshooting and service-restoration ownership.
· Security, compliance and least-privilege mindset.
· Clear communication with global technical and business stakeholders.
· Ability to work across architecture, network, security, application and service-management teams.
· Automation, standardization and continual-improvement orientation.
Skill Requirements
|
Skill Area |
Expected Capability |
Priority |
|
Core AWS Operations |
Accounts, VPC, EC2, EBS, S3, ELB, Auto Scaling, Route 53, backup and recovery. |
Must |
|
Monitoring and Audit |
CloudWatch, CloudTrail, AWS Config, alarms, logs, dashboards and service health. |
Must |
|
Identity and Security |
IAM roles/policies, KMS, Secrets Manager concepts, Security Groups and compliance controls. |
Must |
|
Automation and IaC |
AWS CLI, Python/PowerShell and Terraform or CloudFormation; Git-based change control. |
Good to Have |
|
ITSM / Reliability |
Incident, major incident, problem, change, capacity, lifecycle and service-level management. |
Must |
|
Cost and Governance |
Tagging, inventory, Organizations/SCP awareness, utilization review and optimization recommendations. |
Must |
|
AWS PaaS |
Operational support for RDS, Lambda, SNS/SQS, EKS or other managed services. |
Preferred |
|
DevOps |
GitHub Actions, CodePipeline/CodeBuild or comparable delivery tooling. |
Preferred |
|
Hybrid Operations |
Windows/Linux OS support and hybrid connectivity awareness. |
Preferred |
|
Security Operations |
Security Hub, GuardDuty, Inspector, SOC integration and audit evidence. |
Preferred |