AI Agents for Science
Course description
Modern AI agents combine foundation models with external components and capabilities: tools, memory, planning, feedback, interaction, etc. These systems can search the research literature, write and execute code, operate scientific software, propose experiments, and collaborate with people or other agents. Thus their promise is especially striking in science---but so are the risks of unreliable reasoning, fabricated evidence, irreproducible results, and poorly specified objectives.
This course develops the technical foundations needed to understand and build AI agents, then studies their use across scientific workflows. Students will read current papers, critique systems and evaluations, experiment, and complete a team project.
Learning objectives
By the end of the course, students should be able to:
- explain the core components of modern AI agent systems,
- design agents that use tools, memory, retrieval, and structured feedback,
- evaluate agent behavior with meaningful metrics,
- identify failure modes involving e.g., reasoning, evidence, safety, reproducibility,
- read and critically assess current research on agents for science,
- formulate and carry out a research project in this area.
Prerequisites
Familiarity with basic machine learning is assumed. Students should be comfortable programming and reading research papers. Prior exposure to deep learning or natural language processing is helpful but not strictly required.
Course format
The course has two parts. The first fifteen meetings develop the technical foundations of AI agents: instruction following, tools, environments, planning, memory, frameworks, evaluation, training, safety, and interaction with people and other agents. The remainder of the semester is organized around student-led paper presentations, discussion, and a major research project.
Assessment (tentative)
- Assignments and agent-building exercises: 10%
- Paper presentation and participation: 40%
- Final research project: 50%
Research project
Students will work in small groups on a project that studies, applies, evaluates, or extends AI agents for a scientific task. Projects should make a clear research contribution and include a careful evaluation. Compute and data requirements should be considered early; ambitious ideas with disciplined, reproducible scope are encouraged (but please discuss with me!). Project grading will be based on a group interview/discussion with the instructor.
Academic integrity and responsible use
Coursework must follow UW–Madison academic-integrity policies. Because AI systems are themselves the subject of the class, permitted and required uses of AI tools will be specified for each assignment. Students must disclose tool use, verify generated claims and citations, and remain accountable for submitted work.