System Operations Engineer is responsible for maintaining, monitoring, and optimizing the organization‘s IT infrastructure, servers, networks, and cloud environments to ensure high availability, performance, and security. This role plays a critical part in supporting business continuity, troubleshooting system issues, implementing infrastructure improvements, and driving operational excellence through automation.
Key Responsibilities
1. System Administration & Operations
- Administer and maintain Windows and Linux server environments.
- Monitor system health, performance, and availability using enterprise monitoring tools.
- Perform regular system maintenance, upgrades, and patch management.
- Ensure optimal performance and reliability of business-critical systems.
- Conduct routine system health checks and capacity planning activities.
2. Infrastructure Management
- Manage physical and virtual servers, storage systems, and network infrastructure.
- Deploy, configure, and maintain IT infrastructure solutions.
- Perform data backup, restoration, and disaster recovery activities.
- Support on-premises, hybrid, and cloud-based environments.
- Collaborate with vendors and service providers when required.
3. Monitoring & Incident Management
- Monitor alerts and proactively respond to system incidents.
- Diagnose and resolve infrastructure-related issues within agreed service levels.
- Conduct root cause analysis (RCA) for recurring or major incidents.
- Escalate critical issues and coordinate resolution with relevant teams.
- Participate in on-call support rotations and emergency response activities.
4. Security & Compliance
- Implement and maintain system security controls and best practices.
- Manage user access, permissions, and identity management processes.
- Apply security patches and vulnerability remediation measures.
- Support compliance audits and information security initiatives.
- Ensure adherence to company IT policies and industry standards.
5. Automation & Continuous Improvement
- Develop and maintain automation scripts for operational tasks.
- Improve infrastructure deployment and management processes.
- Support DevOps initiatives and Infrastructure as Code (IaC) practices.
- Identify opportunities to enhance system performance and operational efficiency.
- Recommend and implement technology improvements.
6. Documentation & Knowledge Management
- Create and maintain technical documentation, SOPs, and system diagrams.
- Document incidents, resolutions, and change management activities.
- Maintain accurate records of system configurations and infrastructure assets.
- Contribute to knowledge-sharing initiatives within the IT team.
Qualifications
Education
- Bachelor‘s degree in Information Technology, Computer Science, Information Systems, or a related field.
- Equivalent technical certifications and experience may be considered.
Experience
- At least 5 years of experience in System Administration, IT Operations, Infrastructure Engineering, or a similar role.
- Experience managing enterprise IT infrastructure and production environments.
- Hands-on experience with cloud platforms and virtualization technologies is preferred.
- Experience in incident management and troubleshooting complex technical issues.
Technical Skills
Operating Systems
- Microsoft Windows Server
- Linux (Ubuntu, CentOS, Red Hat Enterprise Linux)
Cloud Platforms
- Microsoft Azure
- Amazon Web Services (AWS)
- Google Cloud Platform (GCP)
Virtualization & Containerization
- VMware vSphere
- Hyper-V
- Docker
- Kubernetes
Monitoring & Logging Tools
- Zabbix
- Prometheus
- Grafana
- ELK Stack
- Splunk
Networking
- TCP/IP
- DNS
- DHCP
- VPN
- Firewall Administration
- Load Balancing
Scripting & Automation
- PowerShell
- Bash/Shell Scripting
- Python
Database Fundamentals
- Microsoft SQL Server
- MySQL
- PostgreSQL
Soft Skills
- Strong analytical and problem-solving abilities.
- Excellent troubleshooting and diagnostic skills.
- Ability to work effectively under pressure and manage multiple priorities.
- Strong communication and collaboration skills.
- Proactive mindset with a focus on continuous improvement.
- High level of accountability and attention to detail.
Preferred Certifications
- Microsoft Certified: Azure Administrator Associate
- AWS Certified SysOps Administrator
- Red Hat Certified System Administrator (RHCSA)
- Red Hat Certified Engineer (RHCE)
- CompTIA Network+
- CompTIA Server+
- ITIL Foundation
- Certified Kubernetes Administrator (CKA)