A 12-week internship on evaluation of long-context models. You will publish an internal report and, if the work holds, an external paper.
What you will do
Reproduce baselines, propose a tighter metric, and run ablations with a mentor from the research team.
What you bring
A current MS or PhD program in ML or a related field, and evidence you can finish an experiment without being managed hour by hour.