IT Operations Engineer

Your day-to-day focus: 

Incident & Problem Management: Own the RCA process for production incidents — diagnose, resolve, and put preventive measures in place so issues don't recur 

Production Monitoring & Support: Continuously monitor service health, detect anomalies early, and act before they become incidents 

Deployment Execution: Design, implement, and maintain CI/CD pipelines using GitHub Actions and related tooling to automate build, test, security scanning, and deployment processes.

Environment Oversight: Keep Pre-Production and Production environments stable and aligned — not building them from scratch, but ensuring they behave as expected day to day 

Runbook & Knowledge Management: Document operational procedures, known issues, and resolution steps to build a reliable knowledge base for the team 

Cross-team Collaboration: Work shoulder-to-shoulder with development and platform teams to triage issues, clarify operational requirements, and close the feedback loop between prod and dev 

 

Operational improvement: 

Identify recurring pain points and propose automation or tooling to reduce toil 

Improve observability coverage — dashboards, alerts, log queries — to catch issues faster 

Contribute to service continuity initiatives and disaster recovery drills 

 

5+ years in IT operations, application support (2nd/3rd line), or a similar production-facing role
Proven track record of owning incidents end-to-end — from alert to RCA to prevention
2+ years working within an ITIL framework (incident, problem, change management)
Experience working in Agile delivery environments alongside development teams
Excellent English communication skills — able to explain technical issues clearly to both engineers and non-technical stakeholders, C1
Must-Have Technical Skills:
Production operations & troubleshooting:
Proficiency with Jenkins to build and maintain pipelines to execute and troubleshoot deployments
Proficiency with CI/CD pipeline – introducing improvements and keeping the pipeline automated
Excellent troubleshooting and problem-solving skills with the ability to independently investigate complex production issues.
Proficiency with log analysis and alerting tools: Splunk, Sysdig
Fluency in observability tooling: Prometheus, Grafana — reading dashboards, tuning alerts
Comfortable operating services running on Kubernetes (checking pod health, reading logs, triggering restarts — not cluster administration)
Excellent skills with Ansible for applying configuration changes in controlled operational scenarios
Strong knowledge of Docker and Docker Compose
Basic scripting skills (Bash, Python) for automation of repetitive operational tasks and reconciliation of data
 
Nice-to-have knowledge and experience:
IBM Datastage operational experience
Awareness or willing to learn Pega, Airflow
 
Application & data layer:
Relational databases (Oracle, DB2) — querying, interpreting execution plans, identifying data-related incidents
Working knowledge of ETL application behavior, Rest API communication
Experience supporting distributed systems and platform services such as Kafka message flow
Java/ development background for understanding the solution and integrations
ID: 3924 job_post.published_on: 20/08/2026
announcement.apply