Natural Language Processing in Customer Service: Transforming Customer Experiences with AI
SOURCE: COMMUNITY.NASSCOM.IN
AUG 06, 2026
Natural Language Processing: How NSF Supports Visionary New Approaches to NLP
SOURCE: CACM.ACM.ORG
JUN 26, 2026
NSF-designed and funded calls for research and training have supported multiple generations of transformative NLP leaders.
By Kathleen McKeown and Christopher D. Manning
Posted Jun 23 2026

Credit: Anastasiia Hevko / Getty Images
Today, large technology firms and AI companies advertise the remarkable commonsense reasoning abilities and contextual understanding of their large language models (LLMs). Less than a decade ago, however, when Yejin Choia was thinking about reviving commonsense reasoning, a seemingly intractable problem that had led to the downfall of symbolic AI nearly 20 years prior, her colleagues warned her that she was being foolish and putting her career at risk. At the time, huge progress had recently been made using artificial neural network (ANN) models in machine learning (ML) for important practical tasks. In 2016, for example, The New York Times ran a prominent feature on how Google Translate had been hugely improved—and simplified—by adopting the new state-of-the-art neural machine translation approach. Still, Choi was not to be deterred. By adopting a large-scale text data perspective, she and her students were able to exploit the first generation of work on LLMs to automatically generate a loosely structured commonsense knowledge base, COMET. She then went on to develop commonsense benchmarks, HellaSwag and Winogrande, that were cited as evaluations in OpenAI’s 2023 announcement of GPT-4. Today, her work is foundational to the most advanced AI tools: Choi’s breakthrough was the result of the funding she had received from the U.S. National Science Foundation (NSF) that allowed her to pursue unpopular—but ultimately transformational—research.
This is precisely the type of work that the NSF has historically enabled: visionary research in fields like natural language processing (NLP) that can yield major impacts on society. Even though it can take years of additional development, often within technology companies, the potential of such discoveries eventually reaches industry or the public. Sometimes breakthroughs come from NSF funding provided to individual faculty members or small team awards, as in Choi’s case. But transformative paradigm shifts within and across disciplines often depend on larger awards, such as the NSF Partnerships for International Research and Education (PIRE) award or the NSF AI Institutes recently established around the country. The NSF’s impact is not limited to technology breakthroughs, of course. As this article will discuss, NSF also funds Ph.D. students, via the NSF Graduate Research Fellowship Program (GRFP), and junior faculty, especially through NSF CAREER awards, who will become the nation’s trailblazing research scientists and industry engineers.
For more than 30 years, NSF funding has been an essential catalyst for the research that made possible the LLMs we see today. As far back as 1990, NSF began funding research on two important lines of research: NLP data collection, and empirical methods. As large quantities of high-quality text data became more accessible, researchers began experimenting with a variety of probabilistic and other ML methods, leading to tremendous improvements in machine translation and to other NLP tools. By the 2010s, neural approaches emerged and in a few more years—once again with NSF support—led to the revolutionary LLMs coming from industry today: models such as OpenAI’s ChatGPT, Anthropic’s Claude, Google’s Gemini, Meta’s Llama, but also Alibaba’s Qwen and Deepseek’s R1.
Just as important, NSF has for decades fostered NLP research creating technology with social impact, as one of the few—if not the only—government agencies expressly championing the principle that the benefits of research should go beyond the academic community and reach society as a whole. This has led to research on technologies that improve access to information globally, address the mental and behavioral health crises in our country, and meet the needs of the most vulnerable, whether the elderly or the disenfranchised.
In this article, we highlight how crucial NSF support has been to both the foundational NLP research that made today’s LLMs possible, as well as applications with important positive social impact. Synergy and balance between these lines of work is essential. For industry to flourish, it needs well-trained young researchers with big ideas and the skills to realize them—and that training is largely supported by the NSF. Conversely, without NSF funding, progress on NLP and LLMs would have been slower to develop and more constrained in its scope, given the comparatively limited and short-term focus of industry. Moreover, there are many ethical issues posed by LLMs and agent-based systems: Will they be accessible to and usable by all segments of the population? Can we ensure the safety of humans as they interact with LLMs? Can LLMs be extended to handle all the languages of the world? Can they help us improve our health, our work, our learning, and our lived experiences? Or will they narrow our opportunities and limit innovation? As corporations drive adoption of LLMs and march toward artificial general intelligence (AGI), academic research unencumbered by corporate goals and profit margins is an indispensable complement to industry activities, providing the accurate and transparent evaluations necessary to ensure we develop LLMs that reflect our broader human and social values.
The path to today’s LLMs can be traced back 30 years to a new approach dubbed “Statistical NLP,” initially pioneered at IBM Research and AT&T Bell Labs, in which the field moved from symbolic, rule-based approaches to empirical, probabilistic, and ML approaches. Adopting ideas from information theory, which was the origin of language modeling, NLP research used these models heavily in text analysis and machine translation. This shift was characterized by greater reliance on large datasets, with a focus on systematically collecting and annotating data, leading to systems based on data and analysis of data. Foreshadowing current trends in NLP, ML demonstrated how we could learn from data to create more robust systems. Instead of prescriptive grammatical approaches, machine learning allowed us to observe how language is actually used, enabling interpretation of the informal language that comprises most of human communication.
The importance of data. The first obstacle to pursuing statistical NLP research was that researchers, whether corporate or academic, did not yet have access to large amounts of digital language data, either text or speech. Thanks to seed funding from both the Defense Advanced Research Projects Agency (DARPA) and NSF, however, this problem was solved by the founding of the Linguistic Data Consortium (LDC) in 1992. Several years of ongoing NSF support sustained the LDC until the development of an independent funding model, based on dues from members and occasional government funding of specific projects. Key to LDC’s origins were Fred Jelinek, then at IBM, who pushed for its establishment to broaden statistical NLP work, and Mark Liberman, a speech researcher who had recently moved from AT&T Bell Labs to the University of Pennsylvania, and who then spent the next 30-plus years building up and developing the LDC. The LDC was instrumental in the shift toward using text data to ground NLP research, enabling empirical NLP research at new companies and in academia. By centralizing contracts with data providers, such as news organizations, the LDC freed individual organizations (both universities and companies) from having to negotiate their own contracts. In contrast, today’s practice of scraping data from the Web without acknowledgment has led to copyright lawsuits (e.g., from The New York Times) that threaten both research and commercial enterprises. A further benefit of the LDC was its careful development of complete and tested annotation guidelines, ensuring data annotations had strong conceptual validity.
The development of statistical NLP: the Johns Hopkins Summer Workshops. Language models that predict which words are likely to appear in a given context were first heavily used in the speech community, starting with Jelinek’s work at IBM. It was Jelinek’s success in using LMs for speech that was the catalyst for applying them to NLP, which, through foundational language-model research funded by the NSF and others, enabled the precursors to today’s NLP technologies.
One particular initiative played a vital role in the development of statistical NLP: the six-week-long Johns Hopkins Summer Workshops. Since 1999, they have been funded by the NSF to convene groups of faculty and students to work on specific hard research problems. The intellectual community and technologies developed in these workshops generate key insights and help disseminate their impact. The workshops have also educated hundreds of students through close research collaboration with faculty and established researchers, many of whom have become leaders in NLP research and major industry contributors to the commercial language-modeling technology used today.
Impact: One of the 1999 workshops focused on the problem of building statistical machine translation (MT) systems. Over the course of a single workshop, participants developed techniques that showed it was possible to train an MT system “in a day.” The legacy of this workshop is illustrated by the later achievements of its participants: Franz Och became the primary creator of Google Translate, while Yaser Al Onaizan managed groups at IBM and Amazon and is currently the deputy CEO of HUMAIN, a leading Saudi AI company. In 2006, a workshop led by Philipp Koehn resulted in the phrase-based, open source Moses machine-translation toolkit. Not only did it become the MT tool most used by academic NLP researchers for the next decade, but its approach formed the basis of early machine-translation systems developed at Google and Microsoft, transforming multilingual NLP worldwide, both inside and outside the academy.
Social impact: Meaningful access to world events. As told by the first author, Kathy McKeown: I got my start with an NSF-funded Presidential Early Career Award, which provided funding for five years of foundational research on language generation. My early foundational work on generation, as well as prior NSF-funded work by other people on interpretation and information extraction, convinced me the time was right for exploring text summarization, with text as both input and output. This was a new, yet important, problem for the field, given the information overload occurring as the Internet grew exponentially. In 1996, my first NSF proposal on text summarization was rejected as “too ambitious.” When NSF launched the Stimulate program a few years later, with an emphasis on funding larger, more ambitious, collaborative projects, my team at Columbia (including Alfred Aho, Shih-Fu Chang, and Judith Klavans) received two grants on summarization that enabled us to look at multimodal summarization of both text and images, as well as statistical techniques for summarization that did not require fully interpreting the input text.
With this initial funding, we focused on summarization of news. We used multiple news articles on the same event to generate a short output paragraph that summarized the consensus across different articles about what had happened. Our initial research branched into Newsblaster (see Figure 1), a summarization platform that showcased our algorithms working on the news of the day. We launched Newsblaster in 2001, one year before the official launch of Google News. It was right after 9/11, and given the terrible events of that day, we knew people would want to follow that event over time. Newsblaster provided an online window into events around the world and allowed tracking of an event over time, identifying new events spurred by the initial one. The platform allowed us to carry out research on many related topics. For example, we developed approaches to generate summary updates that distinguished and highlighted new information from what was already known. We developed a multilingual version of Newsblaster in which the summaries were in English, but drew on—and linked back to—source articles in a variety of languages, enabling a more balanced and holistic view of events.

Figure 1. Newsblaster
Impact: Our work on summarization spurred many follow-on efforts. For example, one student, Dragomir Radev, went on to form large student groups as a faculty member at the University of Michigan (Ann Arbor) and subsequently Yale, to work on summarization. Some of his students likewise went on to form major summarization efforts (e.g., Alex Fabbri in Salesforce). We participated in the formation of the National Institute of Standards and Technology’s Document Understanding Conferences, which provided summarization data and task-based evaluations. This both accelerated research in the field and led to a leaderboard on summarization. Our foundational, NSF-supported work led to an avalanche of interest in the task. Summarization of a news article is today considered a solved problem and is leveraged commercially worldwide on a daily basis. Still, summarization of multiple news documents on the same event remains an active area of research as does summarization of many other genres, such as summarizing legal documents, patient medical records, or narratives.
Just over a decade ago, neural network-based ML models emerged as a key NLP technique that not only increased accuracy on benchmarks but opened new possibilities for learning and representing the semantics and pragmatics of human language use—that is, the meanings of words and phrases, and how they are affected by the linguistic and real-world contexts in which they are used. While this approach has antecedents to the early 2000s (or, indeed, before that), modern deep-learning work in NLP started about 2010 and came to dominate the field around 2014.
Neural methods enable representation of meaning. In 2010, another NLP breakthrough emerged at Johns Hopkins University (JHU) through Tomáš Mikolov’s work in neural language modeling. Funding from an NSF PIRE award—which supports large international group collaborations on foundational research topics—to JHU, Brown University, Charles University in Prague, and Saarland University in Saarbrücken, Germany brought Mikolov, a student at Charles University, together with lead researcher Sanjeev Khudanpur at JHU. The project’s goal was to go beyond syntax (the predominant focus of early work) to explore meaning representations and their applications in language models for speech recognition. Their work set a new state-of-the-art in speech recognition and established beyond any doubt that recurrent neural networks outperformed the then-standard n-gram language models, which had been ubiquitous for the previous 30 years. Concurrently, at Stanford, this article’s second author, Christopher Manning, was approaching human language meaning even more directly by building neural models of meaning composition in phrases and sentences and generated superior benchmark results on paraphrase detection and sentiment analysis.
Impact: These two lines of work led directly to the canonical word2vec and GloVe algorithms for word vectors (continuous meaning representations) in 2013–2014. These discoveries raised the performance of systems on nearly all NLP tasks and introduced the idea of subword modeling that became a stepping stone to early LLMs, such as ELMo, BERT, and GPT in 2017–2018.
Neural machine translation. The first (modern, successful) work in neural machine translation (NMT) was published by researchers at Google in 2014. However, it was the NMT work of Kyunghyun Cho, a postdoc at the University of Montreal, published in 2015, that transformed the field by adding a fundamental new mechanism into neural architectures: attention. Second author Manning’s lab quickly iterated on Cho’s ideas, producing even better results using a different formulation of attention later that same year. Neural attention would go on to revolutionize how neural networks for language and beyond were built, and how scalable and powerful they could be.
In 2010, Cho moved from Korea to Aalto University in Finland because the Academy of Finland, an NSF counterpart in Finland, had agreed to take a chance on his out-of-fashion doctoral research proposal on restricted Boltzmann machines. According to Cho: “If I have managed to make even a tiny bit of contribution to generative AI, it is because the Academy of Finland made a decision to support exploratory and unconventional research, and then I got similar curiosity-driven foundational research support from CIFAR [the Canadian Institute for Advanced Research] and NSERC [the Natural Sciences and Engineering Research Council], in Canada.” CIFAR and NSERC are NSF counterparts in Canada. Cho’s career arc illustrates the importance of a global commitment to supporting and funding pioneering research ideas driven by scientists’ insights and enthusiasm, rather than the necessarily short-sighted objectives of profit-focused companies and even mission-focused government organizations.
Impact: Neural attention was the central innovation that made possible the development of the transformer neural network architecture. This highly scalable transformer architecture quickly led to large, powerful LLMs, and to this day is central to all leading LLMs, such as ChatGPT by OpenAI. It was a key development in the history of AI, the impact of which is still being realized.
Language understanding: Understanding textual inferences. Beyond NMT, the neural NLP work described here by the second author Chris Manning continued to evolve: I had worked for about a decade on machine translation, and neural machine translation was a great setting to develop early end-to-end neural models. The task could be viewed as simply neural sequence transduction by exploiting the large data resources of parallel text (translated text in two languages) that had been built up in the previous decade. As someone with a background in linguistics, however, I wanted to explore whether neural networks could be used for human language understanding and reasoning. Using NSF grants, my colleagues and I conducted initial studies looking at modeling language compositionality as a way to improve sentiment analysis. This task, which decides whether a text expresses positive or negative views about a topic, led to our building the Stanford Sentiment Treebank. I was also keen to try a neural approach to natural language inference (NLI), which focuses on determining the inferential relationships between sentences. For example, given a sentence such as Street patrols deter crime, a model is asked to determine if it “contradicts” (or “entails”) the sentence Street patrols are a catalyst for more crime. In this case, the correct inference is “contradicts.” A neural-network approach demanded data, so we set about crowdsourcing the first large dataset of such examples, the Stanford Natural Language Inference dataset. Though we produced some initial results for compositional neural-network systems on this task, ours were quickly overtaken by others using the new neural attention approaches. This dataset raised awareness of the NLI problem across the neural network researcher community, most of whom otherwise were unaware of it because backgrounds in machine learning, statistics, or even physics were more common than those in linguistics or philosophy.
Impact: NLI became a leading benchmark task for NLP systems, and the lead Ph.D. student on the Stanford NLI dataset team, Sam Bowman, joined the faculty at New York University and produced not only a successor NLI dataset, but the widely used General Language Understanding Evaluation (GLUE) benchmark. The GLUE benchmark, work led by a Ph.D. student funded by an NSF GRFP, combined the benchmarks of various academic researchers, including both the Stanford Sentiment Treebank and the Stanford NLI dataset. Five of its 10 tasks were NLI tasks. For the early pre-trained transformer years from BERT through T5 and ELECTRA (i.e., 2018–2020), GLUE was the standard evaluation of natural language understanding, directing and driving progress in industry labs, until it was replaced in the 2020s by more difficult benchmarks as the abilities of LLMs greatly improved.
Social impact: Mental and behavioral health. In the coming years, the demand for behavioral health professionals in the U.S. is predicted to see a greater than 36% increase—more than triple that of other occupations. This signals an urgent need to address significant training and quality-assessment bottlenecks in one of the fastest-growing workforces in the U.S. At the University of Michigan, Rada Milhacea, a computer science professor, and a cross-disciplinary team of fellow researchers in computer science and health communication are developing NLP tools to support the growing demand for behavioral counseling to address issues like mental health, drug abuse, and weight loss. Through turn-by-turn language analysis (e.g., determining whether a statement is a question or a reflection) and the overall characteristics of a counseling session (e.g., measuring empathy), these researchers are developing methods that can provide automatic, scalable assessment of counseling performance. These approaches—which rely on joint neural and symbolic representations that bring together data-driven insights along with expert knowledge stored in knowledge bases—can provide real-time evaluation and feedback to counselors, as well as offer suggestions to improve the quality of interactions. This recent application of a neural approach illustrates its continued relevance and positive impact on social problems.
Impact: The University of Michigan system has already been used in several public health courses and is being updated for use by counselors in Honduras and Mexico. These methods will help counselors, physicians, and health coaches deliver more effective, personalized care. By enabling ongoing, automated feedback and coaching, the tools aim to improve counseling quality at scale while also advancing a new frontier of NLP research that supports an urgent need for real-world human support.
LLMs rely heavily on foundational research from recent years, supported by NSF, which has been vital to their broader success. For instance, top-p sampling, a sampling technique designed to mitigate output degeneration and ill-formed text, was funded by NSF and is used by all major LLM platforms, both commercial and open source. Flash Attention, an NSF-supported innovation that accelerates the attention mechanism in transformers by two- to three-fold, has become the industry-standard implementation of attention in neural-network libraries.
NSF-sponsored research is also instrumental in accurately evaluating the capability and behavior of modern LLMs, contributing to widely recognized benchmarks such as Massive Multitask Language Understanding (MMLU) and Graduate-Level Google-Proof Q&A (GPQA). MMLU includes roughly 16,000 multiple choice questions that test an LLM’s knowledge across 57 different topics, from STEM to religion. GPQA tests the scientific and reasoning knowledge of LLMs on 448 multiple choice questions written by experts in fields such as biology and chemistry. In addition to testing knowledge, benchmarks also test performance on tasks, such as text summarization, including Multi-News—designed to test performance of summarization on multiple input news documents on the same topic—and SummEval, which was designed to test performance of metrics against human judgments.
Building small LLMs in academia. In 2019, Danqi Chen began her faculty career at Princeton as one of the few people actively trying to build neural language models in academia. Despite not having the compute resources available to even a startup AI company, she was committed to doing this work in academia, not only as a way to give smart young minds the opportunity to contribute to the field but also to ensure both openness and scientific rigor. Work like this lifts the broader innovation ecosystem by providing knowledge on which the whole community can build, rather than there being silos of knowledge within a few huge technology companies. Using relatively cheap structured pruning of an LLM, followed by further training, Chen’s group did early influential work in producing Sheared Llama, small but well-performing 1- and 3-billion-parameter language models. More recently, she introduced new work in instruction-tuning LLMs during post-training, and developed a new, simpler offline preference optimization objective for instruction tuning, SimPO. At launch, her post training of Google’s Gemma-2 model, gemma-2-9b-it-SimPO, was the strongest model with less than 10 billion parameters in the Chatbot Arena evaluation. This work was funded by NSF grants to Chen, both her own CAREER award and another joint research award aimed at developing conceptual understanding of LLMs.
Impact: Small language models are vital to making modern language models available on edge devices such as mobile phones. Chen’s models, such as Sheared Llama 1.3B and Gemma-2-9B-it-SimPO, are widely used; each downloaded more than a million times. However, the larger impact is the influence of the research techniques introduced. Similar techniques to Sheared Llama have been used by Meta and Nvidia, and building small language models has become an increasing emphasis of other big technology companies, including Microsoft, Google, Alibaba, and Hugging Face.
Social impact: Identifying Black individuals at risk. As told by the first author, Kathy McKeown: About 10 years ago, I was approached by social work professor Desmond Patton (now at University of Pennsylvania), who was studying the relationship between the digital and offline lives of Black youth. Over the years, our group, now also joined by a linguist and African American language expert, Jessi Grieser (a professor at University of Michigan), developed techniques to identify online postings expressing emotional distress after traumatic events in Black communities as well as the events that triggered them, enabling valuable follow-up interventions by social workers to provide help. Using LLMs as tools to support Black individuals requires that the tools accurately interpret African American language (AAL). Our evaluations of standard LLMs, however, show they cannot accurately interpret these communications, meaning that mental health and other tools developed using LLMs will fail to meet the needs of many communities. We were working on mitigations to enable better understanding when our NSF grant was abruptly terminated in April 2025 by the federal government.
Impact: The project has the potential for profound impact in society. Automatic flagging of posts expressing emotions of distress would be a unique and important way for social workers and mental health professionals to identify individuals at risk and is far more scalable than their current manual approaches. The creation of LLMs that do understand AAL would also open up other important applications (e.g., in health) to much wider use.
Social impact: Adapting to the world’s languages and people. Truly inclusive and global systems must be able to understand languages other than English, but the LLMs of today typically have a more difficult time with non-English languages and dialects, just as they do with AAL. Thanks to an NSF CAREER Award, Yulia Tsvetkov, a computer science professor at the University of Washington, launched her career with investigations into how well LLMs interpret non-English languages. Her work reveals that the tokenization strategies in commercial LLMs lead to higher costs and poorer performance for many non-English languages, making them far less effective for speakers in lower-income regions and exacerbating global digital inequities. Her team has introduced a large-scale benchmark covering 40 language clusters, 281 dialects, and spanning 10 NLP tasks to address these gaps. Furthermore, she has developed computational methods to analyze biases in Wikipedia biographies, revealing disparities in how people from different gender and racial groups are described across languages. Her grant was also abruptly terminated by the U.S. government in April 2025.
Impact: Tsvetkov’s benchmark enables more inclusive and linguistically robust NLP technologies. Her groundbreaking research on bias—now on hold—is being used to identify and reduce social disparities in widely used resources like Wikipedia content.
AI institutes. In 2019, the NSF launched a call for AI institutes, supporting a new era of multi-institution collaborations that has accelerated U.S. research at the intersection of AI and a range of other disciplines.
At the University of Illinois at Urbana-Champaign, the NSF AI Institute on Molecule Synthesis (MMLI) has assembled an interdisciplinary team of computer scientists and chemists focused on AI for science, resulting in a new Modular Chemical Language Model (mCLM) that trains on interleaved sequences of molecule functional modules and natural language descriptions of these functions (see Figure 2).

Figure 2. Interleaving of molecules and language
Impact: The mCLM model has already improved the chemical functions of 430 FDA-approved drugs on the market, demonstrating how dramatically this model can improve our ability to understand and improve the potential of drugs. Other efforts on AI for Science using LLMs are simultaneously appearing across the U.S., for example at AI2, NYU, and MIT. This will be a major AI research area going forward.
At North Carolina State University, the NSF EngageAI Institute has built a highly interdisciplinary team of education and AI researchers to transform STEM teaching and learning through the development of engaging, AI-driven narrative learning environments.
Impact: By leveraging the power of stories to drive human learning, the Institute’s work will give teachers the tools to create the transformative STEM education environments that will educate and empower the nation’s next generation of STEM leaders.
At Columbia, one thrust of our Institute on Artificial and Natural Intelligence (ARNI) is vision and language, where our interdisciplinary team of cognitive-science, natural-language, and vision researchers is innovating ways to help blind and low-vision people experience art. Currently we are developing a novel multimodal system that can guide users through a personalized, immersive exploration of artwork through a combination of text, speech, and haptics like touch.
Impact: In addition to driving improvements in assistive technology for artwork and other experiences for the 300 million individuals worldwide who are blind or have low vision, this project has generated a new metric for evaluating descriptions of artwork.
Given the incredible capacities of current LLMs, some may believe that further research in NLP is no longer needed, and that remaining challenges will be well addressed by industry. Yet, as we have illustrated, NSF support for fundamental academic research has even recently resulted in advances on which current industry tools now depend. Over more than 30 years, innovation, evaluation, and social impact in NLP has been fostered by NSF-designed and funded calls for research and training, which have supported multiple generations of transformative NLP leaders. Loss of that scientifically driven force would disrupt and delay the next cycle of U.S. advances in the field.
Continued NSF funding would enable research needed to maintain the U.S.’s current fragile leadership in NLP: advances in evaluation, socially impactful AI, methodologies for efficient, small, and safe models, and unimagined applications for NLP, perhaps in scientific research areas.b Unlike industry, academia’s unique environment facilitates interdisciplinary research, enabling the pairing of research in NLP with any discipline across many schools on campus. Research with social impact is also well suited to academia; students find such research meaningful and are energized to participate. Faculty freedom in pursuing research means they can investigate exciting new modeling approaches that typically are not explored in industry. Today, methodologies for evaluation significantly lag LLMs’ apparent capabilities. Without independent, transparent evaluation methods, we will continue to see unexpected, emerging capabilities that may lead to harm. Evaluation also points to avenues for future research. Likewise, exploring socially beneficial applications of AI and NLP often presents the type of genuinely hard problem that may have little commercial appeal at first, but proves revolutionary for the field and society at large. Expanding NLP tools to more language variants not only helps ensure all of society benefits from its technical advances, but can create new markets for these tools. Supporting academic research fosters talent that enables creative approaches to hard technical problems, advancing both our technological capacity and our shared human values. The NSF has begun funding for the second five years of its AI Institutes, and this will spur AI advances, but only if NSF continues to have funding.
Over the course of decades, the two of us have personally had the privilege of working with—and learning from—astonishing students at Columbia and Stanford. These students identify new research problems and develop sophisticated technical approaches, taking research in unexpected directions. Despite our extensive experience, every year these students continue to surprise us. At universities, students in NLP are afforded the unique opportunity to immerse themselves in interdisciplinary collaborations with a wide variety of experts in linguistics, literature, psychology, social science, medicine, machine learning, and computer systems. Our students become faculty at both research universities like MIT, Berkeley, and Princeton as well as at undergraduate-centered institutions like Barnard, Williams, and The Cooper Union. Or they are hired by tech giants such as Google, Microsoft, Anthropic, OpenAI, Amazon, and Meta, critically influencing the directions these companies take as well as their products. As academic leaders at universities and core contributors at leading technology companies, their work has global reach and impact. We have seen this time and time again, with trained researchers going into industry from universities as well as programs such as the Johns Hopkins summer workshops. Industry leaders such as Yaser Al-Onaizan (formerly at Amazon), Sam Bowman (Anthropic), Michel Galley (Microsoft), Sonal Gupta (Microsoft), Franz Och (formerly Google), Fernando Pereira (Google), Dan Roth (Oracle), and Luke Zettlemoyer (Meta), as well as individual research staff members, have benefited from NSF funding at some point in their career.
The creation of this rich talent pool is perhaps the most important product of NSF-funded university research, and is core to the paradigm shifts in NLP that have led to today’s most astonishing technologies.
The recent trend of NSF grant cancellation as well as diminished NSF resources threatens to disrupt these patterns of innovation and the development of NLP talent. While the two canceled grants we mentioned here will undermine the reach of LLMs and their products beyond mainstream English speakers, other grants addressing purely technical NLP topics (e.g., on statistical bias) have also been canceled or indefinitely postponed. Without an adequate number of NSF staff to review proposals, new awards will be dramatically curtailed. Most importantly, what future research innovations and cutting-edge technology companies will never exist because we have cut off the relatively modest public investment that has made them possible up to this point?
Time and time again, we’ve seen academics come up with new ideas for research that were initially deemed “crazy” even by their peers. But these ideas have often led to remarkable new technical advances and tasks for NLP, and they would not have happened for decades—if at all—without the unique funding structure offered by the NSF. We gave a few illustrative examples, but the amount of impactful NLP research that NSF has funded is surprising, leading to the training of thousands of Ph.D. students in the last 10 years alone. While it is hard, in the present moment, to imagine where we might be in 20 years, as long-standing researchers in the field we are confident of one thing: Somewhere during that time, a faculty member or student will come along with a seemingly preposterous research direction that—if provided with the type of intellectual freedom and financial support the NSF has provided over the past three decades—will lead us into another new, exciting, and presently unimaginable future.
This article benefited from the contributions of many individuals who described their NSF-funded efforts. Our thanks goes to Mohit Bansal, Chris Callison-Burch, Danqi Chen, Kyunghyun Cho, Yejin Choi, Jason Eisner, Sanjeev Khudanpur, Heng Ji, Mark Liberman, Rada Mihalcea, and Yulia Tsvetkov for their willingness to contribute by providing advice and actual text. We also thank Susan McGregor, a research scholar at Columbia’s Data Science Institute with a background in journalism, for extensive editing that made our points more forcefully, and Susan Ellington, associate vice president of public affairs at Columbia University, for her critique and suggested edits. We thank Melanie Subbiah, Amith Ananthram, and Yunfan Zhang for comments on an early draft that helped shape our thinking, and the Communications reviewers and editors for their thoughtful comments later on.
FootnotesAbout the Authors
Kathleen R. McKeown is the Henry and Gertrude Rothschild Professor of Computer Science at Columbia University and is also the founding director of the Data Science Institute at Columbia University, New York, NY, USA.
Christopher Manning is the inaugural Thomas M. Siebel Professor in Machine Learning in the Departments of Linguistics and Computer Science at Stanford University and a co-founder and Senior Fellow of the Stanford Institute for Human-Centered Artificial Intelligence (HAI), Stanford, CA, USA.
LATEST NEWS
Machine Learning
Machine Learning Identifies Predictors of Weight Loss in Adolescents With Obesity
AUG 15, 2026
WHAT'S TRENDING
Data Science
5 Imaginative Data Science Projects That Can Make Your Portfolio Stand Out
OCT 05, 2022
SOURCE: COMMUNITY.NASSCOM.IN
AUG 06, 2026
SOURCE: DIGITALJOURNAL.COM
AUG 06, 2026
SOURCE: LEARN.G2.COM
JUL 31, 2026
SOURCE: NEUROSCIENCENEWS.COM
JUL 31, 2026
SOURCE: RESEARCH.GOOGLE
JUL 25, 2026
SOURCE: NERDBOT.COM
JUL 25, 2026
SOURCE: KONICAMINOLTA.COM
JUL 15, 2026
SOURCE: MARBLEHEADCURRENT.ORG
JUL 15, 2026