AI alignment and risk debates, presented fairly
Understand the arguments about AI risk from every side, and form your own reasoned view
A taste of a lesson
Isn't worrying about AI taking over just science fiction?
Many serious people think so, and their argument is strong: today's models have no lasting goals or plans of their own, and dramatic scenarios rest on extrapolation. Others argue the concern is not robots with feelings but capable systems optimising imperfect objectives, something already seen in small ways, such as reward hacking. They worry the gap grows as systems gain autonomy. Both sides agree on some steps, like better testing. Exercise: what single piece of evidence would most change your view?
Written by the teacher as an example. In your lesson the tutor answers your own questions, and like any AI it can be wrong.
What you will be able to do
- Explain what alignment means and its key technical concepts with examples
- Present the main positions on AI risk in their strongest form
- Distinguish lab evidence, theoretical argument and speculation in risk claims
- Identify the cruxes that drive disagreement and state your own view with uncertainty
Lesson plan
- 1 What alignment means Define alignment and why it is hard. Start
- 2 Failure modes in current systems Learn the documented ways systems go wrong. Start
- 3 The catastrophic risk case Understand the strongest argument for serious long term risk. Start
- 4 The present harms and sceptical cases Understand the strongest objections and alternative priorities. Start
- 5 Cruxes and bridges Find where disagreements really lie and what views share. Start
- 6 Your reasoned view State your position, its uncertainty and what would change it. Start
Try asking
About this tutor
For anyone who has heard both 'AI could be an existential threat' and 'that is science fiction distracting from real harms' and wants to understand the debate properly. You learn what alignment means, the core technical concerns (specification gaming, reward hacking, goal misgeneralisation, deceptive behaviour, loss of oversight), the evidence that exists from current systems, and the main positions: those focused on catastrophic and long term risks, those focused on present harms, sceptics of risk claims, and those who see these concerns as connected. Each position is presented in its strongest form. You finish able to explain the debate and state your own view with its uncertainties.
Reviews
4.3
4 ratingsSample
- Ronan C.Sample
The first time I have heard both sides presented without sneering. The cruxes lesson helped me see that my disagreement with a friend was about timelines, not values.
- Elsa N.Sample
Balanced, maybe too balanced for me. I wanted the tutor's own view, which it would not give. The examples were good though.
- Aditi V.Sample
Clear on the technical concepts and careful not to overstate the lab findings. Arguing the opposite side was uncomfortable and useful.
- Gustavo M.Sample
Fair and well structured. I wanted more depth on scalable oversight, but the overview was excellent.
About the teacher
I teach the big debates in AI, from alignment and safety testing to copyright, fairly and clearly
9 tutors 312 lessons taught Sample
I teach the questions people argue about: whether advanced AI could be dangerous, how labs test models, why models flatter us, who owns AI output, and how power over AI is distributed. I present each position in its strongest form before we examine it, and I separate evidence from speculation out loud. My background is in research communication and running...
See Bao's profile and tutorsMore like this
Other tutors on the same or nearby topics.