Good to Great Speaker Series
Good to Great Speaker Series
Organized by Assistant Professor Ju Yeon (Julia) Park
Spring 2027
(Date TBD)
Dr. Adam Berinsky is the Mitsui Professor of Political Science at MIT and serves as the director of the MIT Political Experiments Research Lab (PERL). He is also a Faculty Affiliate at the Institute for Data, Systems, and Society (IDSS). He is the author of three books including “Political Rumors: Why We Accept Misinformation and How to Fight It.” Professional website: https://polisci.mit.edu/people/adam-berinsky
Fall 2026
November 18, 2026, 11:15am-12:30pm
Dr. Arthur Spirling is the Class of 1987 Professor of Politics in the Department of Political Science at Princeton University. His research centers on quantitative methods for analyzing political behavior, especially institutional development and the use of text-as-data. He previously won teaching and mentoring awards at Harvard and NYU, along with the "Emerging Scholar" prize from the Society for Political Methodology. Professional website: https://arthurspirling.org/
September 25, 2026, 11:15am-12:30pm, Derby 2130: In Song Kim, “Measuring Interest Group Positions on Legislation: An AI-Driven Analysis of Lobbying Reports”
Dr. In Song Kim, is an Associate Professor in the Department of Political Science at MIT and a Faculty Affiliate of the Institute for Data, Systems, and Society (IDSS) at Massachusetts Institute of Technology. He specializes in International Political Econom and is well known for his two databases that are widely used in political science: LobbyView and TradeLab. Professional website: https://web.mit.edu/insong/www/
Abstract: Special interest groups (SIGs) in the U.S. participate in a range of political activities, such as lobbying and making campaign donations, to influence policy decisions in the legislative and executive branches. The competing interests of these SIGs have profound implications for global issues such as international trade policies, immigration, climate change, and global health challenges. Despite the significance of understanding SIGs' policy positions, empirical challenges in observing them have often led researchers to rely on indirect measurements or focus on a select few SIGs that publicly support or oppose a limited range of legislation. This study introduces the first large-scale effort to directly measure and predict a wide range of bill positions-Support, Oppose, Engage (Amend and Monitor)- across all legislative bills introduced from the 111th to the 117th Congresses. We leverage an advanced AI framework, including large language models (LLMs) and graph neural networks (GNNs), to develop a scalable pipeline that automatically extracts these positions from lobbying activities, resulting in a dataset of 42k bills annotated with 279k bill positions of 12k SIGs. With this large-scale dataset, we reveal (i) a strong correlation between a bill's progression through legislative process stages and the positions taken by interest groups, (ii) a significant relationship between firm size and lobbying positions, (iii) notable distinctions in lobbying position distribution based on bill subject, and (iv) heterogeneity in the distribution of policy preferences across industries. We introduce a novel framework for examining lobbying strategies and offer opportunities to explore how interest groups shape the political landscape.
Fall 2025
October 14, 2025, 11:15am-12:30pm, Derby 2130: Molly Roberts, “Propaganda is already influencing large language models: evidence from training data, audits, and real-world usage”
Dr. Margaret (Molly) Roberts, is a Professor in the Department of Political Science at UC San Diego. She is a co-direct the China Data Lab at the 21st Century China Center and an affiliate at the UC Institute on Global Conflict and Cooperation. Professional website: https://margaretroberts.net
Abstract: Millions of people around the world query (prompt) large language models for information. While several studies have compellingly documented the persuasive potential of these models, there is limited evidence of who or what influences the models themselves, leading to a flurry of concerns about which companies and governments build and regulate the models. We show through six studies that coordinated propaganda from powerful global political institutions already influences the output of U.S.-based large language models via their training data. We first provide evidence that Chinese state propaganda appears in large language model training datasets. To evaluate the plausible effect of this inclusion, we use an open-weight model to show that additional pre-training on Chinese state propaganda generates more positive answers to prompts about Chinese political institutions and leaders. We link this phenomenon to commercial models through two audit studies demonstrating that prompting models in Chinese generates more positive responses about China's institutions and leaders than the same queries in English. We then use a cross-national audit study to show that languages of countries with lower media freedom exhibit a stronger pro-regime valence than those with higher media freedom. The combination of influence and persuasive potential suggest the troubling conclusion that states and powerful institutions have increased strategic incentives to disseminate propaganda in the hopes of shaping model behavior.
Spring 2025
April 14, 2025, 11:15am-12:30pm, Derby 2130: Bruce A. Desmarais, "Public Officials’ Online Sharing of Low-factual Content: Institutional and Ideological Checks"
Bruce A. Desmarais (he/him), is the DeGrandis-McCourtney Early Career Professor in Political Science, Director of the Center for Social Data Analytics, and Co-Hire for the Institute for Computational and Data Sciences at Pennsylvania State University. Professional Website: brucedesmarais.com
Abstract: Elected officials occupy privileged positions in public communication about important topics—roles that extend to the digital world. In the same way that public officials stand to lead constructive online dialogue, they also hold the potential to accelerate the dissemination of harmful content. We explore and explain the sharing of misinformation, which we refer to as low-factual content, by examining nearly 500,000 Facebook posts by U.S. state legislators from 2020 to 2021. We validate a widely used low-factual content detection approach in misinformation studies, and apply the measure to all of the posts we collect. Our findings reveal that the prevalence is relatively rare, affecting less than one percent of legislators’ posts overall. However, Republican legislators share low-factual content at higher rates, and certain states emerge as hotspots for such content. We also find that conservative lawmakers are more likely to share such content, with this tendency potentially intensifying in conservative districts, and waning in liberal ones. Legislative professionalism plays a significant role in misinformation circulation: legislators with higher institutional capacity and resources tend to be less inclined to share low-factual information, which implies that heightened professional standards may help reduce misinformation. We conclude with a discussion of the implications of our findings for future interventions to reduce the spread of low-factual content.
Fall 2024
October 2, 2024, 11:15am-12:30pm, Derby 2130: Naoki Egami, "Using Large Language Model Annotations for the Social Sciences: A General Framework of Using Predicted Variables in Statistical Analyses"
Dr. Naoki Egami, an Assistant Professor at the Department of Political Science at Columbia University, gave us a talk on how to use Large Language Models for a data annotation task while avoiding any bias for the downstream analysis. Email: naoki.egami@Columbia.Edu. Professional Website: naokiegami.com
Abstract: Social scientists use automated annotation methods, such as supervised machine learn-ing and, more recently, large language models (LLMs), that can predict labels and generate text-based variables. While such predicted text-based variables are often analyzed as if they were observed without errors, we show that ignoring prediction errors in the automated annotation step leads to substantial bias and invalid confidence intervals in downstream analyses, even if the accuracy of the automated annotations is high, e.g., above 90%. We propose a framework of design-based supervised learning (DSL) that can provide valid statistical estimates, even when predicted variables contain non-random pre- diction errors. DSL employs a doubly robust procedure to combine predicted labels and a smaller number of expert annotations. DSL allows scholars to apply advances in LLMs to social science research while maintaining statistical validity. We illustrate its general applicability using two applications where the outcome and independent variables are text-based.
Spring 2024
April 16, 2024: Michelle Torres, "Beyond Prediction: Identifying Latent Treatments in Images"
Michelle Torres, an assistant professor at UCLA, is a political methodologist with expertise in image analysis using computer vision and machine learning techniques. Her research classifies political visual messages/frames to understand their role in the generation and processing of political information.
Abstract: Images are a rich and crucial element of political communication. The complexity of the information they convey creates challenges for the identification, interpretation, and explanation of the effects of visual messages on information processing and attitude formation. In this article, we adapt a methodological approach used in text analysis, the supervised Indian Buffet Process (sIBP) developed by Fong and Grimmer (F&G, 2016, 2021), to identify latent treatments in images and evaluate their impact on outcomes of interest. First, we use a convolutional neural network (CNN) to decompose images into substantively meaningful and interpretable tokens, visual words, to then form the input of the sIBP. Then, we follow the framework introduced by F&G and demonstrate the utility of this approach using two datasets: 1) a novel experiment measuring attitudes towards climate change in response to visual frames and 2) images of the Black Lives Matter (BLM) movement protests manually labeled by human coders according to the level of conflict they depict. We find significant differences between demographic groups in the way they perceive images, and also unmask latent treatments that confound the relationship between our treatment and outcomes of interest. Importantly, this paper extends the usage of computer vision tools in social sciences beyond prediction of image labels to uncovering, understanding, and visualizing the features of images that produce outcomes.