Skip to main content
Talent Arabia logo

AI Ops Engineer (AI FinOps, Governance, Reliability, and Production Support)

Talent Arabia
17 hours ago
Contract
On-site
Abu Dhabi, United Arab Emirates
Automation

Urgent requirement for AI Ops Engineer (AI FinOps, Governance, Reliability, and Production Support) in banking domain is required for our banking clients in Abu Dhabi ,UAE
Design and manage standardized CI/CD pipelines, release workflows, deployment automation, promotion controls, and governance for AI applications, agents, and platform services.--MustImplement operational controls for models and AI assets, ensuring versioning, traceability, compliance, auditability, and safe AI releases across environments.--MustEnable canary deployments, controlled rollouts, rollback strategies, AI quality evaluations, monitoring, telemetry, dashboards, runbooks, and production-readiness practices.--MustDrive AI FinOps, cost visibility, token/model usage monitoring, capacity management, self-service templates, operational playbooks, and enterprise-wide standards for scalable AI operations.--Must


Banking Domain --Must
 Role PurposeWe are seeking an AI Ops Engineer to establish the operational backbone for enterprise AI platforms, enabling application teams to release AI products safely, repeatedly, and at scale. The role is accountable for production release discipline, LLMOps practices, deployment automation, operational governance, cost visibility, and self-service operating standards for AI-native delivery teams.Key Responsibilities

  • AI Release Engineering & CI/CD: build and design standard release pipelines and promotion controls for AI applications, agents, platform services, and configuration changes across environments, ensuring repeatable deployment, governance, and release evidence.
  • LLMOps & AI Lifecycle Management: embed operating controls for models, AI assetsand release for auditable.
  • Progressive Delivery: Implement deployment patterns to reduce production risk, including controlled rollout, canary release, and rollback readiness
  • Evaluation, Observability & Production Readiness: Embed AI quality checks, operational telemetry, dashboards, runbooks, and readiness criteria into the delivery lifecycle so AI services are measurable and supportable.
  • AI FinOps & Capacity Governance: Provide visibility and controls for AI workload consumption, including model usage, token spend, platform capacity, quota management, and optimization opportunities.
  • Self-Service & Continuous Improvement: Convert proven operating patterns into reusable templates, release standards, onboarding guidance, operational playbooks, and paved-road workflows that allow teams to move quickly while maintaining enterprise control.
  • CI/CD & Release Engineering (GitHub Actions or similar)
  • LLMOps / MLOps / AI Lifecycle Management
  • Cloud-native Platform Operations & Observability
  • AI FinOps, Governance, Reliability, and Production Support
  • Cross-functional collaboration with Platform, SRE, QA, Security, Architecture, and Product teams.

Required Experience

  • Strong production engineering background operating cloud-native, AI, or high-scale API platforms in an enterprise environment.
  • Hands-on experience with CI/CD, GitHub Actions or equivalent automation, and environment management.
  • Working knowledge of LLMOps, telemetry, release governance, and production-readiness practices.
  • Experience with observability, change control, service reliability, and continuous operational improvement.

Ability to partner with platform engineering, QA/SRE, cybersecurity, architecture, product, and delivery teams to standardize safe and scalable AI operations