Senior ML Infrastructure Engineer
Quick answer
Ellison Institute of Technology is hiring a remote-friendly Senior ML Infrastructure Engineer (Oxford, England, United Kingdom) with a competitive salary and benefits to build and operate high-performance ML compute clusters.
- Role
- Senior ML Infrastructure Engineer
- Organization
- Ellison Institute of Technology
- Location
- Oxford, England, United Kingdom (remote friendly)
- Work setup
- Remote
- Level
- Senior
- Compensation
- Competitive salary (dependent on experience) + travel allowance + bonus
- Category
- Engineering & Technology
- Apply by
- 2026-12-31
The role
Join our SciComp team to build the cloud and compute foundation that enables scientific breakthroughs. Deliver reliable, secure platforms and self-service guardrails that accelerate experimentation and turn ideas into results - faster, at scale, and with confidence.
What you'll do
- Build, operate, and continuously optimise our high-performance GPU training and inference clusters
- Drive systems design and implementation for high-throughput data paths
- Proactively benchmark, profile, and resolve performance bottlenecks
- Establish comprehensive observability, resilience, and automated security controls
- Partner with Research, Data, and Applied teams to forecast capacity and cost for GPU and storage needs
What it takes
- Proven experience leading the design, build, and operation of high-performance ML compute clusters at scale
- Expertise with high-throughput storage systems for ML/HPC workloads
- Expert-level understanding of GPU architecture, high-speed networking for distributed training, and performance profiling to resolve bottlenecks
- A proactive, autonomous approach to systems design and the proven ability and desire to ideate, co-create and implement optimal solutions
- Exposure to migrating or transforming ML infrastructure from traditional schedulers to modern, containerised systems
- Expert-level understanding of IaC and CI/CD practices (e.g., Terraform, Argo CD)
What you'll bring
How we treat you
Enhanced holiday, Pension - Employer contribution 7.5%, minimum employee contribution 5%, Life Assurance, Income Protection, Private Medical Insurance as standard, Employee discounts, Electric car scheme, Nursery Salary Sacrifice scheme, Cycle to Work Scheme, Family Planning, Neurodiversity support, Coaching & Therapy services
Frequently asked questions
Where is the job located?
The job is located in Oxford, England, United Kingdom, but it is remote-friendly.
What is the compensation?
The compensation is a competitive salary (dependent on experience) + travel allowance + bonus.
What are the key qualifications required?
Key qualifications include proven experience leading the design, build, and operation of high-performance ML compute clusters at scale, expertise with high-throughput storage systems for ML/HPC workloads, and expert-level understanding of GPU architecture, high-speed networking for distributed training, and performance profiling to resolve bottlenecks.
How do I apply?
You can apply now by visiting the Ellison Institute of Technology's job posting page.
How to apply
Apply directly on Ellison Institute of Technology's site. We link straight through — no resume parsing, no profile to fill out.
This listing is aggregated from a third-party source and its summary may be auto-generated, so details can be inaccurate or out of date. ForGood is not the employer and is not liable for the content — please verify everything on Ellison Institute of Technology's official posting before applying.