How to apply for CUDA Engineer

Fuse Energy

About Fuse Energy

Fuse Energy is a fully integrated energy company that develops its own solar and battery projects, builds hardware, trades power in real time, and sells directly to consumers. It has raised over $200M from investors including Balderton, Lakestar, Accel, and Creandum. The company is now expanding into high-performance compute infrastructure at the intersection of energy and AI, which is why it needs GPU kernel engineers.

About the role

You will write and optimise custom CUDA kernels for transformer inference operations, tuning performance across memory bandwidth and compute bottlenecks. The work happens at the level of SMs, warps, and memory hierarchies, with the goal of squeezing maximum throughput out of every GPU in the fleet. This role directly supports Fuse's expansion into data centre compute, which is one of the fastest-growing sources of electricity demand.

A typical day

A typical day involves profiling existing inference kernels, writing or modifying CUDA code for a specific transformer operation, and testing performance across different GPU configurations. You will spend time in Nsight, reading PTX or SASS, and discussing trade-offs with the team about memory layout, fusion, and throughput targets. The exact rhythm depends on the team, so ask about sprint structure and deployment cadence in interviews.

Who Fuse Energy is looking for

  • Strong CUDA C/C++ experience with hands-on kernel development, not just using high-level frameworks like PyTorch or TensorFlow.
  • Deep understanding of GPU architecture: SMs, warps, shared memory, registers, occupancy, and memory coalescing.
  • Experience optimising transformer inference workloads, including attention, matrix multiplication, and layer normalisation kernels.
  • Ability to profile and debug GPU code using tools like Nsight Compute or Nsight Systems, and to reason about roofline models and arithmetic intensity.

Tips for this application

  • Lead with concrete CUDA work: kernel names, speedups achieved, GPU models used, and profiling data. Vague claims about 'GPU optimisation' will not pass review.
  • Mention any experience with transformer inference specifically. If you have written attention or GEMM kernels, describe the tiling, vectorisation, or memory layout decisions you made.
  • Show that you understand why an energy company is hiring CUDA engineers. Connect your work to reducing cost or power per inference token.
  • Apply directly through Fuse Energy's careers page if one exists, or via the listed job board. Avoid sending generic applications to multiple roles at the same company.
  • If you have open-source CUDA contributions, include links. This is one of the few fields where public kernel code is a strong signal.

What to cover in your cover letter

Focus on specific CUDA kernels you have written and optimised, with numbers: latency reduction, throughput gain, or memory bandwidth utilisation. Explain how you approach a transformer inference operation from first principles, including how you decide between fusing operations, changing memory layout, or increasing occupancy. Show awareness that Fuse is an energy company using compute to lower costs, not a pure AI lab. Mention any experience working with production inference systems where reliability and power efficiency matter.

Draft a cover letter

Research before applying

  • Read Fuse Energy's public materials on their integrated model: solar, batteries, trading, and now compute. Understand why they are building data centre infrastructure.
  • Look up the investors named in the job description and any public statements they have made about Fuse's compute plans.
  • Research typical transformer inference bottlenecks and recent CUDA optimisation techniques for attention and GEMM, so you can speak concretely in interviews.
  • Check whether Fuse has published any engineering blog posts or talks about their GPU work. If not, prepare questions about the current state of their inference stack.

Likely interview topics

Based on the job description, expect questions about:

  • Write or reason about a CUDA kernel for a transformer operation such as softmax, layer norm, or scaled dot-product attention.
  • Explain how you would diagnose a kernel that is memory-bandwidth bound versus compute bound, and what changes you would make in each case.
  • Discuss trade-offs between shared memory usage, register pressure, and occupancy on a specific GPU architecture.
  • Describe how you would profile an end-to-end inference workload and identify the top bottleneck.
  • Ask about the GPU fleet, the inference stack, and how kernel work integrates with the rest of the energy and trading systems.
Practise interview questions

Common mistakes to avoid

  • Listing CUDA as a skill without describing any kernel you have actually written or optimised. This role requires proof of low-level work.
  • Focusing only on model training or high-level ML frameworks. The job is about inference kernels and hardware-level performance.
  • Ignoring the energy context. Fuse is not a generic AI company. Candidates who do not connect GPU efficiency to energy cost or grid demand miss the point of the role.

Deadline

No deadline is listed. Roles without a deadline usually close once the employer has enough candidates, so apply soon if you are interested.