AI Inference Engineer
Fuse Energy
Posted
Jul 23, 2026
Location
Remote
Type
Full-time
Mission
What you will drive
- Define Fuse's inference serving strategy and architecture from first principles.
- Design and build the serving stack: request routing, batching, scheduling, and autoscaling for high-throughput, latency-sensitive inference workloads.
- Own model-level optimisation strategy for serving, including quantisation, distillation, speculative decoding, and similar techniques.
- Act as a direct technical owner of inference performance and reliability, working closely with CUDA and GPU engineering teams.
Impact
The difference you'll make
This role enables Fuse to deliver high-performance AI inference at scale, pairing real power delivery with real compute to accelerate renewable energy adoption and reduce the carbon footprint of AI workloads.
Profile
What makes you a great fit
- 4+ years of experience building or operating large-scale inference serving systems.
- Deep, hands-on experience with inference serving frameworks and optimisation techniques (batching, KV-cache management, quantisation, speculative decoding).
- Strong systems thinking and ability to reason about the full path from incoming request to served response across a large cluster.
- Comfortable working directly with GPU/CUDA engineers and making high-stakes architecture calls.
Benefits
What's in it for you
- Competitive salary and an equity sign-on bonus.
- Biannual bonus scheme.
- Fully expensed tech to match your needs.
- Breakfast and dinner allowance for office based employees.
About
Inside Fuse Energy
Fuse Energy is a forward-thinking renewable energy startup on a mission to deliver a terawatt of renewable energy fast, combining first-principles thinking with cutting-edge technology to build a radically better energy system.