Summary
Keywords
Full Transcript
October 30, 2024 Joseph Jay Williams, University of Toronto Learn more about the speaker: https://www.psych.utoronto.ca/people/directories/all-faculty/joseph-jay-williams This lecture is from Stanford CS329H: Machine Learning from Human Preferences Machine learning from human preferences investigates mechanisms for capturing human and societal preferences and values in artificial intelligence (AI) systems and applications, e.g., for socio-technical applications such as algorithmic fairness and many language and robotics tasks when reward functions are otherwise challenging to specify quantitatively. While learning from human preferences has emerged as an increasingly important component of modern AI, e.g., credited with advancing the state of the art in language modeling and reinforcement learning, existing approaches are largely reinvented independently in each subfield, with limited connections drawn among them. This course will cover the foundations of learning from human preferences from first principles and outline connections to the growing literature on the topic. This includes but is not limited to: -Inverse reinforcement learning, which uses human preferences to specify the reinforcement learning reward function -Metric elicitation, which uses human preferences to specify tradeoffs for cost-sensitive classification -Reinforcement learning from human feedback, where human preferences are used to align a pre-trained language model View the course website: https://web.stanford.edu/class/cs329h/index.html Enroll in the course: https://online.stanford.edu/courses/cs329h-machine-learning-human-preferences
