How to apply for Machine Learning Engineer

Elicit

About Elicit

Elicit is a public benefit company building an AI research assistant to help researchers make better decisions. The product is designed as a scalable ML system that prioritizes systematicity and transparency, with supervision of process rather than outcomes. Working there means contributing to a tool that aims to improve how research conclusions are extracted and used.

About the role

As a machine learning engineer, you will compose tens to thousands of calls to language models to accomplish tasks that a single call cannot. You will curate datasets for fine-tuning, set up evaluation metrics to judge model improvements, and scale semantic search from a few thousand to 100k documents. This role directly shapes the core ML system that powers Elicit's research assistant.

A typical day

A typical day might involve writing code to chain language model calls, reviewing dataset samples for fine-tuning, and checking evaluation metrics on recent model changes. You could also work on scaling the semantic search index and discussing with the team what counts as a process improvement. Since the role is remote, most collaboration happens through written updates and video calls.

Who Elicit is looking for

  • Has hands-on machine learning engineering experience, especially with language models and AI systems in production.
  • Can design and curate datasets for fine-tuning models, for example to extract policy conclusions from papers.
  • Has set up evaluation metrics that determine what changes to models or training setups count as improvements.
  • Has worked with semantic search systems and scaled them from small document sets to 100k or more.

Tips for this application

  • In your resume or cover letter, describe a specific project where you composed multiple calls to language models to solve a task. Include the number of calls, the task, and the result.
  • Explain how you curated a dataset for fine-tuning. Mention the source of data, labeling process, and how you measured quality.
  • Describe an evaluation metric you built for an ML system. State what it measured, how you validated it, and how it changed model decisions.
  • Show experience scaling a semantic search system. Give the starting and ending document counts, the retrieval method, and the main bottleneck you fixed.
  • Reference Elicit's public benefit status and its focus on supervision of process, not outcomes. Connect that to your own approach to building transparent ML systems.

What to cover in your cover letter

Focus on your experience composing multiple language model calls into a single workflow. Highlight a dataset you curated for fine-tuning and the evaluation metrics you designed. Emphasize any semantic search scaling you have done, especially from thousands to 100k documents. Mention why Elicit's public benefit structure and process supervision appeal to you.

Draft a cover letter

Research before applying

  • Use Elicit's product to understand how it extracts conclusions from papers and where language model calls are composed.
  • Read Elicit's public benefit company status and any public statements on supervision of process, not outcomes.
  • Look for Elicit engineering blog posts or talks about their ML system, semantic search, or evaluation methods.
  • Check the Elicit website for current research areas or case studies that show what tasks the AI research assistant handles.
Elicit website

Likely interview topics

Based on the job description, expect questions about:

  • How you would compose tens to thousands of language model calls to extract policy conclusions from research papers.
  • Your process for curating a fine-tuning dataset, including sourcing, labeling, and quality checks.
  • How you design evaluation metrics that determine whether a model change is an improvement.
  • Your approach to scaling semantic search from a few thousand to 100k documents, including indexing and retrieval trade-offs.
  • How you think about transparency and systematicity in an ML system, given Elicit's supervision of process, not outcomes.
Practise interview questions

Common mistakes to avoid

  • Do not describe only single-call language model projects. This role requires composing many calls into a larger workflow.
  • Do not claim experience with semantic search scaling without giving concrete document counts or retrieval methods.
  • Do not treat evaluation metrics as an afterthought. Elicit explicitly requires setting up metrics that define improvements.

Deadline

No deadline is listed. Roles without a deadline usually close once the employer has enough candidates, so apply soon if you are interested.