MLOps Engineer
Keter AI is a senior-led AI consulting and engineering boutique that takes organisations from AI ambition to working systems, from boardroom to codebase.
As an MLOps Engineer you deploy and run models where our clients need them: in their own cloud accounts, in private data centres and in air-gapped environments. You take systems from prototype to production and keep them healthy there, covering Kubernetes and GPU scheduling, model serving with vLLM, observability, evaluation pipelines, CI/CD, infrastructure as code and security. The role is open to mid and senior engineers and is fully remote.
What we look for
At least three years in MLOps, platform engineering or SRE, including models or LLM services you took to production and kept running.
Kubernetes in depth, on managed services and on clusters you built yourself, including GPU scheduling, capacity planning and sharing GPUs between workloads.
Model serving with vLLM or a comparable inference server such as NVIDIA Triton or KServe, with safe rollouts (canary, blue/green) and a tested rollback.
Observability that answers real questions: latency, throughput, error rates, GPU utilisation and cost per request, using Prometheus, Grafana, OpenTelemetry or similar.
Evaluation built into delivery: quality and regression checks for models and LLM applications that run in CI/CD, and drift monitoring once a system is live.
Infrastructure as code and CI/CD with Terraform, Helm and GitOps on at least one of AWS, Azure or GCP, and comfort with on-prem constraints such as no outbound internet access.
Security as part of the work, not an afterthought: secrets management, network policies, least-privilege access and audit logs that a client's CISO can review.
Clear written and spoken English, and the habit of documenting decisions so a client's own engineers can run what you hand over. Polish is an advantage.
What we offer
EUR 5,500-7,800 (about USD 6,200-8,700) per month, depending on experience, on a B2B contract (net, plus VAT where applicable).
Fully remote, with working hours anchored in European time zones. Contractors based in Poland are welcome.
A senior team: you work directly with the engineers and consultants who scope, build and sign off each engagement, with no layers in between.
Real client work: production systems in client clouds and on-prem environments, across sovereign AI, custom AI systems and AI-assisted engineering.
Ownership from architecture to handover, and a say in the tools and patterns we standardise on.
A yearly learning budget for courses, certifications and conference time, agreed with you.
The selection process
Introductory call. Your background, the systems you have run and what you want next.
Technical conversation. With our engineers, built around a deployment you have run and a realistic client scenario: architecture, failure modes and trade-offs.
Final conversation. Scope, rate and start date, followed by a written offer.
Before you apply
Send a short note through our contact page with links to work you have delivered, such as a repository, a write-up or a talk. Tell us your time zone and your earliest start date.