Women in AI Research (WiAIR)

Women in AI Research (WiAIR)

WiAIR
Држава Сједињене Државе
Жанрови Образовање
Језик EN
Епизоде 32
Последња 16.09.2026

Women in AI Research (WiAIR) is a podcast dedicated to celebrating the remarkable contributions of female AI researchers from around the globe. It challenges the perception that AI research is predominantly male-driven and aims to empower early career researchers, especially women, to pursue their passion for AI. Listeners learn from women at different career stages, stay updated on the latest research and advancements, and hear powerful stories of overcoming obstacles and breaking stereotypes.

Епизоде

  • What Makes a Sentence Memorable? Inside Language, Memory and LLMs, with Dr. Greta Tuckute 16.09.2026 1ч 16мин
    Why don't bigger LLMs look more like the human brain? In the brain's language network, today's large models explain only slightly more variance than GPT-2 XL - and the reason says a lot about what those brain regions actually do.Dr. Greta Tuckute (Research Fellow at the Kempner Institute for the Study of Natural and Artificial Intelligence, Harvard) joins Jekaterina Novikova on Women in AI Research to unpack what brain-LLM alignment does and does not tell us. Greta works where neuroscience, cognitive science and AI meet - and when asked to keep only one of those three labels, AI is the first one she drops.The conversation covers how brain alignment develops over training, how sparse autoencoders can turn the "you are just comparing one black box to another" critique into something testable, and what memory experiments reveal about how meaning is stored - from why "pineapple" sticks in memory and "light" does not, to which sentence embeddings predict what people remember.In this episode:Brain alignment tracks formal linguistic competence, not reasoning - and it emerges after not much more than a developmentally plausible ~100M tokens, not 300BWhy better next-word prediction stops meaning "more brain-like" once a model has mastered languageSparse autoencoder features plus surprisal: for one frontal brain region, surprisal alone does almost as well as 32,000 SAE features, while some voxels are captured by just six features- Brains and LLMs share the main, high-variance features of language - not the idiosyncratic onesWhy learning from BPE tokens puts brain-LLM comparisons on "pretty shaky ground", and what should come nextA distinctive meaning makes words and sentences memorable - and SBERT predicts human sentence memory better than the other embedding models testedFrom running a photography business at 14 to a PhD at MIT, and why trying the other path first was worth itPAPERS DISCUSSEDFrom Language to Cognition: How LLMs Outgrow the Human Language NetworkInterpreting Brain Responses to Language with Sparse Features from Language ModelsDriving and suppressing the human language network using large language models Intrinsically memorable words have unique associations with their meaningsA distinctive meaning makes a sentence memorableMENTIONED IN THIS EPISODEAnna Ivanova on formal vs functional linguistic competence (WiAIR)GRETA TUCKUTEWebsite: http://www.tuckute.comBluesky: https://bsky.app/profile/gretatuckute.bsky.socialX: https://x.com/GretaTuckuteWomen in AI Research (WiAIR) is a podcast and YouTube channel where Jekaterina Novikova talks with women doing AI research about their work, the questions driving it, and the paths that brought them there.🎧 Subscribe to stay updated on new episodes spotlighting brilliant women shaping the future of AI.Follow WiAIR at:⁠⁠⁠LinkedIn⁠⁠⁠⁠⁠⁠Bluesky⁠⁠⁠⁠⁠⁠X (Twitter)⁠⁠⁠⁠⁠⁠⁠WiAIR website⁠
  • Is an Image Worth a Thousand Words? Hidden Failures of Multimodal Metrics, with Dr. Elisa Kreiss 19.08.2026 1ч 17мин
    An image is not worth a thousand words - it's worth an indefinite number of them. So why do the metrics we use to evaluate AI-generated image descriptions still assume there's one correct answer?In this episode of Women in AI Research, I talk with Elisa Kreiss (Assistant Professor of Communication at UCLA, director of the Coalas Lab) about what happens when you actually test the metrics the field relies on, and why CLIPScore, one of the most widely used measures for scoring image descriptions, stops correlating with human judgment the moment you introduce context. We also get into why longer descriptions aren't necessarily more informative, what happens when you just ask a model to "be concise," and whether AI models trip over charts and graphs the same way humans do.Elisa's research sits at the intersection of linguistics, accessibility, and multimodal AI, and this conversation covers the full arc of her work, from the theoretical question of why humans never describe images the same way twice, to the practical question of what that means for building systems that actually work for blind and low-vision users.In this episode:Why "context matters" is more radical than it sounds for image description evaluationThe hidden reason CLIPScore breaks down once context enters the pictureWhy length is a bad proxy for information density - and what to use insteadWhat happens when you prompt a model to just "be concise"Why charts and photos need completely different evaluation approachesWhether AI models make the same mistakes as humans when reading data visualizationsWhat NeurIPS's Top Reviewer Award taught her about writing a genuinely useful peer reviewResources & Links:Context Matters for Image Descriptions for Accessibility: Challenges for Referenceless Evaluation MetricsWhen More Words Say Less: Decoupling Length and Specificity in Image Description EvaluationCHART-6: Human-Centered Evaluation of Data Visualization Understanding in Vision-Language Models🎧 Subscribe to stay updated on new episodes spotlighting brilliant women shaping the future of AI.Follow WiAIR at:⁠⁠LinkedIn⁠⁠⁠⁠Bluesky⁠⁠⁠⁠X (Twitter)⁠⁠⁠⁠⁠WiAIR website⁠
  • Is Your AI Just Flattering You? Sycophancy, AI Policies, and More, with Dr. Malihe Alikhani 22.07.2026 1ч 7мин
    Only 19% of Americans say AI has actually improved their productivity - so why the gap between the hype and reality? In this episode of Women in AI Research, Dr. Malihe Alikhani (Northeastern University, Contextual AI Lab) unpacks the hidden failures in how we build and deploy AI: why sycophancy is really a collapse of alignment, why bigger models aren't better aligned, and why "thin" alignment breaks down in the real world.Key topicsThe impact of moving across different AI contexts on system designThe role of language as performative and active in shaping realityInteractive inference and uncertainty in AI systemsThe importance of context in meaning and system designAI policy, transparency, and societal impactSycophantic behavior in large language modelsMeasuring AI alignment: thin vs. thickAI adoption across sectors and demographic groupsThe role of policy in AI development and safetyEthical considerations in AI research and deploymentResources & Links:Breaking the AI Mirror: Sycophancy, productivity, and the future of collaborationHype and harm: Why we must ask harder questions about AI and its alignment with human valuesHow are Americans using AI? Evidence from a nationwide surveyConnect with Dr. Malihe Alikhani:https://x.com/malihealikhani🎧 Subscribe to stay updated on new episodes spotlighting brilliant women shaping the future of AI.Follow WiAIR at:⁠LinkedIn⁠⁠Bluesky⁠⁠X (Twitter)⁠⁠⁠WiAIR website⁠
  • Can Language Alone Create Intelligence? Insights from Neuroscience and AI, with Dr. Anna Ivanova 17.06.2026 1ч 11мин
    Do large language models truly understand language—or are they sophisticated pattern matchers?In this conversation, Dr. Anna Ivanova (Asst. Prof. at Georgia Tech) explores one of the important questions in AI: the relationship between language, thought, and intelligence. Drawing from neuroscience, cognitive science, and AI research, Anna explains why language understanding is harder to define than most people realize, why reasoning and language are not the same thing, and what today's LLMs can and cannot tell us about human cognition.Key Topics:Do LLMs understand language or merely generate convincing text?The difference between formal and functional linguistic competenceWhat LLMs can learn from language alone—and what they cannotWhy human cognition and AI cognition may be fundamentally differentTheory of mind, reasoning, and common misconceptions about AI capabilitiesHow cognitive scientists evaluate the "thinking" abilities of LLMsWhat neuroscience can teach AI researchers about interpretabilityWhy understanding AI requires studying both behavior and internal representationsThe future of multimodal models and AI cognitionResources & Links:What does it mean to understand language?Dissociating language and thought in large language modelsHow to evaluate the cognitive abilities of LLMsHow Do LLMs Use Their Depth?True LensConnect with Dr. Anna Ivanova:https://bsky.app/profile/neuranna.bsky.socialhttps://x.com/neuranna🎧 Subscribe to stay updated on new episodes spotlighting brilliant women shaping the future of AI.Follow WiAIR at:LinkedInBlueskyX (Twitter)⁠WiAIR website⁠
  • 100% Jailbreak Success? The Hard Truth About AI Safety, with Dr. Saadia Gabriel (Part 2) 17.04.2026 33мин
    What actually happens when AI systems fail in the real world?In this final part of our conversation with Saadia Gabriel (UCLA), we unpack one of the most urgent challenges in modern AI: why even the most advanced models remain vulnerable to manipulation - and what that means for safety, fairness, and society.From multi-turn jailbreaking attacks with near 100% success rates to misinformation shaping human beliefs, this conversation goes beyond surface-level concerns and dives into how harms actually emerge in deployed systems.We explore:Why current guardrails are not enoughHow realistic attack scenarios differ from academic benchmarksThe connection between model vulnerabilities and societal harmWhat AI can (and cannot) do about misinformation and persuasionThe open research problems that still don’t have solutionsResources & Links:Generative AI in the Era of 'Alternative Facts'ModelCitizens: Representing Community Voices in Online SafetyTranslation as a Scalable Proxy for Multilingual EvaluationConnect with Dr. Saadia Gabriel:https://x.com/GabrielSaadiahttps://bsky.app/profile/skgabrie.bsky.social
  • From Hate Speech to Best Paper: Building Safer AI Systems, with Dr. Saadia Gabriel (Part 1) 15.04.2026 29мин
    What does it mean to build AI systems we can actually trust?In this first part of our conversation with Saadia Gabriel (UCLA), we explore the deeply personal and technical journey behind her work on AI safety, misuse, and responsible NLP.From experiencing targeted hate speech firsthand to receiving a best paper nomination, Saadia shares how her lived experience shaped her research — and why language models must be designed with both capability and risk in mind.🧠 In this episode, we cover:How personal experiences influence AI research directionsThe intersection of NLP, security, and privacyWhy LLMs can be both powerful and dangerousWhat it means to build trustworthy AI systemsLessons from working across multiple research paradigmsHow to pursue high-impact research as a PhD or early-career scientistResources & Links:X-Teaming: Multi-Turn Jailbreaks and Defenses with Adaptive Multi-AgentsConnect with Dr. Saadia Gabriel:https://x.com/GabrielSaadiahttps://bsky.app/profile/skgabrie.bsky.social
  • EACL 2026: LLMs Can Hear… But Can They Reason? A New Benchmark for Audio Intelligence 13.04.2026 18мин
    What does it actually mean for a model to understand audioPaper: https://arxiv.org/abs/2601.19673In this episode, I talk with Iwona Christop, a PhD student at Adam Mickiewicz University, about her recent EACL paper introducing ART (Audio Reasoning Tasks) — a new benchmark designed to evaluate whether multimodal LLMs can truly reason over audio, not just transcribe or classify it.Most existing benchmarks test audio skills in isolation (like ASR or classification). But real-world intelligence requires something deeper: combining signals, comparing sounds, tracking context, and making decisions.This work takes a different approach:No text-only shortcuts — tasks can’t be solved via transcription aloneReasoning-first design — models must combine multiple audio cuesNo expert knowledge required — anyone can verify correctnessWe also dive into the diverse task design, including:Audio arithmetic (counting and comparing sounds)Cross-recording speaker & language identificationSound-based reasoning (e.g., inferring properties from audio)Speech feature comparison (accents, variations)Multimodal reasoning across text and soundThe dataset includes 9 tasks, 9,000 samples, and 30+ hours of audio — all generated in a scalable way using templates and TTS.👉 If you care about multimodal reasoning, evaluation, or the limits of current LLM capabilities, this conversation is for you.Iwona Christop:https://www.linkedin.com/in/iwona-christop/👍 Like & subscribe for more deep dives into cutting-edge AI research🔔 New episodes from EACL 2026 coming soon#WiAIR #EACL2026
  • EACL 2026: LLMs Can Call Tools -- But Can They Understand Them? 12.04.2026 22мин
    LLM-based agents are everywhere, but most research focuses on just one step: getting the model to call the right tool. What happens after that?Paper: https://arxiv.org/abs/2510.15955In this talk, Kiran Kate (IBM Research) presents new findings from their EACL 2026 paper on a largely overlooked problem:👉 Can LLMs actually understand and use the outputs returned by tools?As tool-augmented systems become more complex, this question becomes critical. The work dives into how current models handle non-trivial, real-world tool responses, and where they break down.💡 Key ideas covered:Why tool calling is only half the story in LLM agentsThe challenge of processing complex tool outputsFailure modes in current LLM-based systemsWhat this means for building robust, real-world AI agentsThis talk is especially relevant if you're working on:LLM agents and tool useEvaluation of LLM capabilitiesReal-world deployment of AI systemsAgentic workflows and reasoning pipelinesKiran Kate: https://www.linkedin.com/in/kiran-kate-8b98672/👍 Like & subscribe for more deep dives into cutting-edge AI research🔔 New episodes from EACL 2026 coming soon
  • EACL 2026: Reasoning Can Hurt LLM Safety?! Rethinking Accuracy in AI Systems 10.04.2026 21мин
    In this episode of #WiAIRpodcast, we dive into a subtle but critical question: Does adding reasoning actually make LLMs safer and more reliable?Paper: https://arxiv.org/abs/2510.21049Atoosa Chegini (University of Maryland, Apple) presents Reasoning's Razor (EACL 2026), where she and her collaborators examine how reasoning impacts high-stakes binary classification tasks, including safety filtering and hallucination detection.Their findings highlight an important nuance:While reasoning can improve overall accuracy, it may degrade performance at low false positive rates -- exactly where real-world systems need to operate.This conversation covers:Why accuracy is a misleading metric for safety-critical LLM applicationsThe importance of evaluating models at fixed false positive rates (FPR)How two models with identical accuracy can behave completely differently in deploymentThe impact of "think-on" (with reasoning) vs "think-off" (no reasoning) settingsPractical implications for RLHF, SFT, and post-training pipelinesIf you're working on:LLM evaluation & reliabilityAI safety or hallucination detectionProduction deployment of language models— this discussion offers a perspective that is both technically grounded and immediately actionable.Atoosa:https://www.linkedin.com/in/atoosa-chegini-6713741a3/https://scholar.google.com/citations?user=5nY9tagAAAAJ&hl=en&oi=ao👍 Like & subscribe for more deep dives into cutting-edge AI research🔔 New episodes from EACL 2026 coming soon
  • EACL 2026: Why LLMs Hallucinate, and How to Make Them Say "I Don't Know" 03.04.2026 15мин
    LLMs are notoriously overconfident, but can we teach them to admit uncertainty?In this episode, Maor Juliet Lavi (Tel Aviv University) presents her EACL 2026 paper on Detecting Unanswerability in Large Language Models with Linear Directions.Paper: https://arxiv.org/abs/2509.22449We cover:Why prompt-based fixes for hallucinations aren’t enoughHow “unanswerability” emerges inside model representationsA simple but powerful idea: linear directions in hidden statesWhy this method generalizes better across datasetsWhat this reveals about where abstract concepts live inside LLMs👉 Watch to see how far we can push models toward knowing when they don’t know.Maor:https://www.linkedin.com/in/maor-juliet-lavi-07494a155👍 Like & subscribe for more deep dives into cutting-edge AI research🔔 New episodes from EACL 2026 coming soon
  • EACL 2026: From Paraphrases to Diagnostics: A Fine-Grained Framework for LLM Auditing 01.04.2026 16мин
    LLMs often give different answers to the same question, just phrased differently. But how do we measure and understand this behaviour rigorously?In this episode of the #WiAIRpodcast, Cléa Chataigner (Mila, McGill) presents AUGMENT, a user-grounded, controlled paraphrasing framework for auditing prompt sensitivity in large language models, accepted as an oral at EACL 2026.Paper: https://arxiv.org/abs/2505.03563Instead of relying on noisy, unconstrained paraphrasing, AUGMENT introduces:Structured, linguistically grounded paraphrase types (e.g., voice, style, dialect)A guided generation + automated quality control pipelineFine-grained analysis of how specific linguistic variations impact model behaviourKey insights:Different paraphrase types can shift model performance in opposite directionsStandard baselines can hide critical failure modesEven strong LLMs struggle with certain transformations (e.g., voice changes)Evaluated across:Bias benchmarks (BBQ)Knowledge tasks (MMLU)Multiple open-source model familiesCléa:https://scholar.google.com/citations?user=NdToDmMAAAA👍 Like & subscribe for more deep dives into cutting-edge AI research🔔 New episodes from EACL 2026 coming soon
  • EACL 2026: You're Using Persona Prompting Wrong 30.03.2026 15мин
    How much control do persona prompts actually give us over LLM behaviour?In this episode of #WiAIRpodcast, Jing Yang (TU Berlin) speaks about the study on persona prompting in socially sensitive tasks, including hate speech detection, sentiment analysis, and commonsense reasoning.Paper: https://arxiv.org/abs/2601.20757The paper takes a closer look at a common assumption: that adding demographic or identity-based personas can help align model outputs with different user groups.In this conversation, we discuss:Whether persona prompting meaningfully changes model predictionsWhy simulated personas don’t necessarily align with real-world demographicsThe gap between improving labels vs. improving model rationalesEvidence that LLMs may systematically over-predict harmful contentWhat this means for synthetic data generation and evaluation practicesOne of the key takeaways is that persona prompting has limited effect as a steering mechanism in these settings ,and should be applied with care, especially in high-stakes or socially sensitive applications.Jing Yang:https://www.linkedin.com/in/jing-yang-7b07aa135👍 Like & subscribe for more deep dives into cutting-edge AI research🔔 New episodes from EACL 2026 coming soon
  • EACL 2026: Why Your LLM Needs Math to Think Better 27.03.2026 19мин
    What if adding math data actually improves reasoning in non-math tasks?In this episode of #WiAIRpodcast, Syeda from the Carnegie Mellon University and NVIDIA presents Nemotron CrossThink (EACL 2026 Oral) - a framework that challenges how we train LLMs for reasoning.Paper: https://arxiv.org/abs/2504.13941Key insights you don’t want to miss:Why multi-domain training beats specialized modelsThe surprising role of math in improving general reasoningHow to get +performance with ~40% less dataWhy harder training samples are better than more dataHow LLMs learn adaptive reasoning strategies across tasksEven more interesting:CrossThink models use 28% fewer tokens → cheaper inferenceThey outperform math-only models on reasoning tasksResults generalize across model familiesSyeda:https://x.com/__SyedaAkterhttps://bsky.app/profile/reasyaay.bsky.socialhttps://www.linkedin.com/in/syeda-nahida-akter-989770114/👍 Like & subscribe for more deep dives into cutting-edge AI research🔔 New episodes from EACL 2026 coming soon
  • EACL 2026: Inference-Time Steering Is Riskier Than You Think 25.03.2026 10мин
    Are the techniques we use to control language models quietly making them less safe?Paper: https://arxiv.org/abs/2602.06256In this EACL 2026 paper, Navita Goyal (University of Maryland) challenges a core assumption behind inference-time interventions: that we can precisely steer model behaviour without unintended side effects.▶️ Key insight:While steering methods improve target behaviours (like reducing over-refusal), they significantly weaken robustness — especially under adversarial prompts and jailbreaks.⚠️ The result?Models become more vulnerable exactly where safety matters most.We also discuss:Why current steering methods lack true specificityThe overlooked gap between in-distribution vs. robust evaluationEvidence of systematic vulnerabilities introduced by steeringA new perspective on disentangling model representationsNavita Goyal:https://x.com/navitagoyal_https://bsky.app/profile/navitagoyal.bsky.socialhttps://navitagoyal.github.io/👍 Like & subscribe for more deep dives into cutting-edge AI research🔔 New episodes from EACL 2026 coming soon
  • EACL 2026: Do LLMs Spread Misinformation Better Than Humans? 23.03.2026 15мин
    What happens when humans and large language models try to persuade each other?In this episode of Women in AI Research, we sit down with Angana Borah (University of Michigan) to unpack her EACL 2026 paper on misinformation, persuasion, and demographic effects in human–LLM interactions.Paper: https://arxiv.org/abs/2503.02038Key insights from the paper:LLM-generated persuasion can reduce human accuracy in identifying true vs false informationSome demographic groups are more susceptible to misinformation than othersHuman-written persuasion can actually improve LLM reasoning (in some cases)LLM agents can form echo chambers, mirroring human group dynamicsHomogeneous groups → higher risk of misinformation propagationEven more surprising:LLMs often score higher than humans on linguistic persuasion metrics — raising serious questions about their role in shaping beliefs.We also discuss:Can LLMs simulate human susceptibility to misinformationWhat happens in multi-turn persuasion scenarios?Where this research is heading next—and what it means for AI safetyAngana Borah:https://x.com/AnganaBorah2https://blueskydirectory.com/profiles/anganaborah.bsky.sociahttps://www.linkedin.com/in/anganaborah/👍 Like & subscribe for more deep dives into cutting-edge AI research🔔 New episodes from EACL 2026 coming soon
  • Does Liking Yellow Make You a School Bus Driver? Hidden Failures in LLMs, with Dr. Hila Gonen 04.03.2026 59мин
    In this conversation, Dr. Hila Gonen (Assistant Professor at the University of British Columbia) joins us to explore the deep insights into how large language models (LLMs) leak semantic information, behave across languages, and how researchers can uncover their root causes. Dr. Gonen shares her journey in interpreting AI systems, addressing biases, and controlling model outputs for safer, fairer applications.In this episode:The influence of prompt elements, like colour, on model predictionsHow semantic leakage impacts model outputs unintentionallyThe role of multilinguality and modality in model safety and behaviourInterventional vs. observational approaches to understanding modelsChallenges in controlling and aligning AI behavior across languages and domainsFuture directions in model interpretability, safety, and causal analysisKey Topics:Color and semantic influence on language model completionsThe concept of semantic leakage and examples from real promptsDifferences between bias, hallucination, and leakage failuresUnintended behaviours discovered through experimentationThe importance of model interpretability and transparencyRoots of behaviour: training data and internal representationsInterventional analysis as a causal tool in NLP researchCross-lingual and cross-modal alignment in safety detectionChallenges in evaluating safety across languages and modalitiesStrategies for building robust controls against unseen attack typesThe future of AI research: combining performance with reliability and safetyEthical considerations: avoiding directions that hinder societal benefitsResources & Links:Lipstick on a Pig: Debiasing Methods Cover up Systematic Gender Biases in Word Embeddings But do not Remove ThemDoes Liking Yellow Imply Driving a School Bus? Semantic Leakage in Language ModelsRewriting History: A Recipe for Interventional Analyses to Study Data Effects on Model BehaviorOMNIGUARD: An Efficient Approach for AI Safety Moderation Across Modalities and LanguagesConnect with Dr. Hila Gonen:LinkedInhttps://x.com/hila_gonenNote: This episode emphasizes practical and theoretical challenges in model interpretability, safety, bias detection, and causality—providing a comprehensive view suitable for researchers, practitioners, and AI enthusiasts interested in responsible AI development.🎧 Subscribe to stay updated on new episodes spotlighting brilliant women shaping the future of AI.⁠⁠WiAIR website⁠⁠Follow us at:⁠⁠LinkedIn⁠⁠⁠⁠Bluesky⁠⁠⁠⁠X (Twitter)
  • Faithfulness and Hallucinations in Reasoning Models, with Dr. Letitia Parcalabescu 11.02.2026 1ч 3мин
    Are reasoning models actually reasoning — or just producing convincing stories?Our guest in this episode of #WiAIRpodcast is Letitia Parcalabescu, the creator of the  @AICoffeeBreak  youtube channel. Letitia joins Jekaterina Novikova for a deep dive into the topics of faithfulness, self-consistency, hallucinations, and the reliability illusion in LLMs and multimodal reasoning models.We discuss why chain-of-thought explanations may not reflect what the model actually did, why RAG does not automatically fix hallucinations, and how vision–language models often rely far more on text than images. We also explore new approaches for grounding and rejection — and why models struggle to say "I don't know."Instead of focusing only on benchmark scores, this conversation asks: What kind of evidence do we need to truly trust reasoning models?REFERENCES:On Measuring Faithfulness or Self-consistency of Natural Language ExplanationsDo Vision & Language Decoders use Images and Text equally? How Self-consistent are their Explanations?Bounding Hallucinations: Information-Theoretic Guarantees for RAG Systems via Merlin-Arthur ProtocolsAI Coffee Break with Letitiahttps://www.youtube.com/c/AICoffeeBreakhttps://x.com/AICoffeeBreak🎧 Subscribe to stay updated on new episodes spotlighting brilliant women shaping the future of AI.⁠WiAIR website⁠Follow us at:⁠LinkedIn⁠⁠Bluesky⁠⁠X (Twitter)
  • AI Safety Beyond Benchmarks -- Dr. Swabha Swayamdipta on Evaluation, Personalization, and Control 21.01.2026 1ч 1мин
    As language models become more capable, the hardest questions are no longer just about performance, but about safety, interpretation, and control.In this episode of Women in AI Research, we speak with Swabha Swayamdipta, Assistant Professor of Computer Science at the University of Southern California and co-Associate Director of the USC Center for AI and Society. Swabha’s research examines how the design and deployment of language models intersect with real-world risks — from how models behave in unexpected ways to how seemingly technical choices can have broader societal consequences.We talk about AI safety from multiple angles: what it means when hidden inputs to models can sometimes be inferred from their outputs, why personalization introduces new trade-offs around privacy and user agency, and how assumptions about model behavior can quietly shape downstream harms. Rather than focusing only on accuracy or benchmarks, the conversation asks what kinds of evidence we actually need to trust these systems in practice.REFERENCESBetter Language Model Inversion by Compactly Representing Next-Token DistributionsImproving Language Model Personas via Rationalization with Psychological ScaffoldsOATH-Frames: Characterizing Online Attitudes Towards Homelessness with LLM AssistantsUncovering Intervention Opportunities for Suicide Prevention with Language Model Assistants🎧 Subscribe to stay updated on new episodes spotlighting brilliant women shaping the future of AI.⁠WiAIR website⁠Follow us at:⁠LinkedIn⁠⁠Bluesky⁠⁠X (Twitter)
  • Do LLMs Understand Meaning? Neuroscience, Evaluation, and the Future of AI, with Dr. Maria Ryskina 31.12.2025 1ч 3мин
    Do large language models actually understand meaning — or are we over-interpreting impressive behavior?In this episode, we speak with Dr. Maria Ryskina, CIFAR AI Safety Postdoctoral Fellow at the Vector Institute for AI, whose research bridges neuroscience, cognitive science, and artificial intelligence. Together, we unpack what the brain can (and cannot) teach us about modern AI systems — and why current evaluation paradigms may be missing something fundamental.We explore how language models can predict brain activity in regions linked to visual processing, what this reveals about cross-modal knowledge, and why scale alone may not resolve deeper conceptual gaps in AI. The conversation also tackles the growing importance of interpretability, especially as AI systems become more embedded in high-stakes, real-world contexts.Beyond technical questions, Maria shares why community matters in AI research, particularly for underrepresented groups — and how diversity directly shapes the kinds of scientific questions we ask and the systems we ultimately build.REFERENCESGender Shades: Intersectional Accuracy Disparities in Commercial Gender ClassificationStereotypes and Smut: The (Mis)representation of Non-cisgender Identities by Text-to-Image ModelsLanguage models align with brain regions that represent concepts across modalitiesElements of World Knowledge (EWoK): A Cognition-Inspired Framework for Evaluating Basic World Knowledge in Language ModelsPrompting is not a substitute for probability measurements in large language modelsAuxiliary task demands mask the capabilities of smaller language models🎧 Subscribe to stay updated on new episodes spotlighting brilliant women shaping the future of AI.WiAIR websiteFollow us at:LinkedInBlueskyX (Twitter)
  • How Does AI Reflect Society, with Dr. Maria Antoniak 10.12.2025 1ч 11мин
    AI doesn’t just process text — it takes in our cultures, reflects our hierarchies, and can make existing power structures even stronger. In this episode of Women in AI Research, Jekaterina Novikova and Malikeh Ehgaghi speak with Dr. Maria Antoniak (Assistant Professor at the University of Colorado Boulder) about inclusivity in AI, the dynamics of cultural representation, what trust in AI really means, why LLMs tend to homogenize research cultures, and what maternal healthcare reveals about the deepest ethical challenges in this field.REFERENCES:⁠Trust No Bot⁠⁠A Large-Scale Analysis of Public-Facing, Community-Built Chatbots on Character.AI⁠⁠LLMs as Rsearch Tools⁠⁠Culture is Not Trivia⁠⁠Research Borderlands⁠⁠NLP for Maternal Healthcare⁠⁠Data Feminism⁠⁠Epistemic Diversity and Knowledge Collapse in Large Language Models⁠⁠Empire of AI⁠🎧 Subscribe to stay updated on new episodes spotlighting brilliant women shaping the future of AI.⁠⁠⁠⁠WiAIR website⁠⁠⁠⁠Follow us at:⁠⁠⁠⁠LinkedIn⁠⁠⁠⁠⁠⁠⁠⁠Bluesky⁠⁠⁠⁠⁠⁠⁠⁠X (Twitter)⁠

Популаран у

Овај подкаст се појављује и у подкаст листама ових земаља.