This one is closed
Live roles like this one
- B 1d ago
-
B
1d ago
Intelligent Infrastructure Engineer
Bright Vision Technologies United States $100k - $150k/yr
-
B
2d ago
Rackspace US, Inc. United States $166k - $243k/yr
-
B
1w ago
Director, Full- stack Engineer, Shopping Tech (Remote- Eligible)
Capital One United States $245k - $279k/yr
See every "HPC Infrastructure Engineer" role →
Get new “HPC Infrastructure Engineer” roles by email
One email a day with what is new in "HPC Infrastructure Engineer". Nothing new, no email.
We confirm the address first, and every mail carries an unsubscribe link. Alerts are ours, not a third party's.
Why this grade This listing scored 48/100, which is a D. It lost the most ground on pay transparency. See the breakdown
- Description depth 20 / 20 How much the posting actually says about the work, measured in characters of real text.
- Pay transparency 12 / 25 A published salary range, worth more than any other single factor because it is what a candidate cannot find out without applying.
- Freshness 8 / 15 How recently it was posted. Older postings are likelier to be filled or abandoned.
- Remote clarity 8 / 15 Whether "remote" means anywhere, or is quietly restricted to one country.
- Corroboration 5 / 10 Whether more than one source carries this listing.
- Role specificity 0 / 10 Whether the listing is tagged well enough to tell what the role actually is.
-5 Ghost-job penalty — Deducted for signals that this posting may not be a real, currently-open role — staleness, repeated relisting, or talent-pool language.
Every figure above is arithmetic over the posting itself — its salary field, its text, its age, its tags and how many sources carry it. How the grades work →
About ElevenLabs
ElevenLabs is an AI research and product company transforming how we interact with technology.
We launched in January 2023 with the first human-like AI voice model. Today, we serve millions of users and thousands of businesses - from fast-growing startups to large enterprises like Deutsche Telekom and Meta. Our investors are some of the world's most prominent, including Andreessen Horowitz, ICONIQ Growth and Sequoia. We've raised $781M in funding and our last valuation was $22B - multiples of 11, always.
We have expanded from voice into three main platforms:
ElevenAgents enables businesses to deliver seamless and intelligent customer experiences, with the integrations, testing, monitoring, and reliability necessary to deploy voice and chat agents at scale.
ElevenCreative empowers creators and marketers to generate and edit speech, music, image, and video across 70+ languages.
ElevenAPI gives developers access to our leading AI audio foundational models.
Everything we do is the result of the creativity and commitment of our team - builders doing the best work of their lives. We are researchers, engineers, and operators. IOI medalists and ex-founders. If you want to work hard and create lasting positive impact, we want to hear from you.
How we work
High-velocity: Rapid experimentation, lean autonomous teams, and minimal bureaucracy.
Impact not job titles: We don’t have job titles. Instead, it’s about the impact you have. No task is above or beneath you.
AI first: We use AI to move faster with higher-quality results. We do this across the whole company—from engineering to growth to operations.
Excellence everywhere: Everything we do should match the quality of our AI models.
Global team: We prioritize your talent, not your location.
What we offer
Innovative culture: You’ll be part of a generational opportunity to define the trajectory of AI, surrounded by a team pushing the boundaries of what’s possible.
Growth paths: Joining ElevenLabs means joining a dynamic team with countless opportunities to drive impact - beyond your immediate role and responsibilities.
Learning & development: ElevenLabs proactively supports professional development through an annual discretionary stipend.
Social travel: We also provide an annual discretionary stipend to meet up with colleagues each year, however you choose.
Annual company offsite: Each year, we bring the entire team together in a new location - past offsites have included Croatia and Italy.
Co-working: If you’re not located near one of our main hubs, we offer a monthly co-working stipend.
About the role
Every model we train runs on infrastructure this role owns. We operate NVIDIA GPU clusters across bare metal and rented capacity, and we're looking for an engineer to join our small research infrastructure team and make that compute fast, reliable, and boring - in the best sense. When the clusters just work, research moves faster. Your impact is measured directly in training throughput and researcher velocity.
This is a builder-operator role with real breadth: one week you're writing automation that eliminates a whole class of manual work, the next you're benchmarking a new provider's InfiniBand fabric or on-site bringing new hardware online. You'll have unusual scope and autonomy - we're a lean team where decisions are made by the people closest to the problem.
What you’ll be doing
Operate and improve our GPU fleet end to end: provisioning, scheduling, monitoring, upgrades, capacity planning
Build automation that keeps the fleet healthy without human intervention — node health checks, automated draining and remediation, burn-in pipelines for new capacity
Own the stack beneath the training code: OS images, NVIDIA drivers, CUDA, container runtimes, NCCL, high-speed networking (InfiniBand/RoCE)
Run and tune job scheduling (Slurm or similar) so researchers get compute fairly and fast
Build and maintain high-performance storage for datasets and checkpoints
Hunt down performance problems: stragglers, degraded links, thermal issues, flaky GPUs — and fix the class of problem, not just the instance
Evaluate rented GPU capacity: benchmark it, validate it, hold providers to their SLAs
Hands-on hardware work when it's needed: racking, cabling, diagnostics, coordinating with datacenter staff and vendors
Keep clusters secure by default: access control, network isolation, secrets
Requirements
Have run large-scale Linux server or GPU environments in production and enjoy both building and operating
Know the NVIDIA stack well — drivers, CUDA, NCCL, DCGM — or have deep systems experience and learn hardware stacks fast
Are comfortable with bare-metal environments, server hardware, and high-speed networking
Write solid automation in Python and/or Bash, with IaC tools like Ansible or Terraform
Are happy digging into noisy data (metrics, logs, PromQL) to find what's actually wrong
Like owning real scope end to end and being the person others rely on
Don't consider any task above or beneath you — datacenter trips included
Nice to have
Experience supporting ML training workloads from the infra side (distributed training failure modes, checkpointing patterns)
Experience evaluating and working with GPU cloud providers
Parallel filesystems (WEKA, VAST, etc) or large-scale object storage
BMC/IPMI/Redfish automation, PXE provisioning at scale
Power and cooling awareness for dense GPU deployments
Location
This role is remote and can be executed globally. If you prefer, you can work from our offices in London, New York, San Francisco, and Warsaw.
#LI-Remote
#HPC-Infrastructure-Engineer-GPU-Clusters
We are an equal opportunity employer and do not discriminate on the basis of race, religion, national origin, gender, sexual orientation, age, veteran status, disability or other legally protected statuses.
Apply for this role Opens jobs.ashbyhq.com — verified as the employer's own application page
Quick question · anonymous · one tap
Would you apply to this job?
Answer to see what other job seekers said.
Your turn · no account needed
Help the next applicant
You may know something about this listing that we cannot see from here. One tap. No account needed. Signed-in reports earn points once the evidence agrees with you.
I know what it pays
Sign in with Google to earn points for reports — 100 confirmed points buy a week of Early Access.
Where this listing came from
- 03 Sep 2026 Company ATS employer ATS first sighting
Seen on 1 board over 35 days. The employer edited the description 2× since we first recorded it.