Technical AI Safety Progam
The Technical AI Safety Program is a 8-week course which covers topics like:
- developing safer AI through RLHF (reinforcement learning with from human feedback), scalable oversight, etc
- evaluating AI risks and safety
- interpreting LLMs (large language models)
- safety research by OpenAI, DeepMind, and Anthropic
We base the seminar on BlueDot Impact's Technical AI Safety course and the ARENA curriculum.
How it works:
Depending of the background of applicants, the program may be in 2 different formats:
- Reading & discussion based: each week, you will do about 1 hour of independent reading (including videos and articles). Then, we will have a 1-hour meeting discussing the material.
- Coding based: spanning the 8 weeks, we will learn how to build neural networks, LLM interpretability tools, reinforcement learning, etc, with less frequent meetings.
The seminar will last 8 weeks, beginning on the week of September 14th and ending on the week of November 9th (skipping the Fall break week). This is an estimate, as we are both testing a new curriculum this semester and experimenting with a more free-flowing discussion.
Apply
Applications are due by 09/06 11:59PM EST.