Technical AI Safety Progam

The Technical AI Safety Program is a 8-week course which covers topics like:

  1. developing safer AI through RLHF (reinforcement learning with from human feedback), scalable oversight, etc
  2. evaluating AI risks and safety
  3. interpreting LLMs (large language models)
  4. safety research by OpenAI, DeepMind, and Anthropic

We base the seminar on BlueDot Impact's Technical AI Safety course and the ARENA curriculum.

How it works: Depending of the background of applicants, the program may be in 2 different formats:

  1. Reading & discussion based: each week, you will do about 1 hour of independent reading (including videos and articles). Then, we will have a 1-hour meeting discussing the material.
  2. Coding based: spanning the 8 weeks, we will learn how to build neural networks, LLM interpretability tools, reinforcement learning, etc, with less frequent meetings.

The seminar will last 8 weeks, beginning on the week of September 14th and ending on the week of November 9th (skipping the Fall break week). This is an estimate, as we are both testing a new curriculum this semester and experimenting with a more free-flowing discussion.

Apply

Applications are due by 09/06 11:59PM EST.