U of A Researchers Help Build AI-Powered Platform to Transform Interdisciplinary Scientific Discovery
With $4.6 million in National Science Foundation funding, researchers are developing an open-source AI platform designed to make scientific data easier to find, connect and use across disciplines.
Photo courtesy College of Science.
Scientific discovery has a data problem.
Across fields ranging from astronomy and biology to agriculture, engineering and public health, researchers are generating enormous amounts of data. But finding the right data, understanding how it is organized and combining it with information collected by other research teams can take weeks or even months—time that could otherwise be spent on analysis and discovery.
A new interdisciplinary project involving researchers across the University of New Mexico, University of Arizona and University of North Carolina at Chapel Hill aims to change that.
Barney Maccabe, Professor and Associate Dean of Research, College of Information Science.
The researchers have been awarded $4.6 million from the National Science Foundation to develop the Multidisciplinary Environment for Scientific Advancement, or MESA, an open-source platform that uses agentic artificial intelligence to help researchers find, organize, connect and use scientific data across disciplines. The U of A will receive $2.1 million over two years for its portion of the project. Tyson Swetnam, associate professor of computer science at UNM, leads the overall project, while Barney Maccabe, professor and associate dean of research in the College of Information Science, serves as U of A’s lead principal investigator.
Terrell Russell serves as co-PI representing UNC’s Renaissance Computing Institute (RENCI), where he serves as director of data management.
Agentic AI Enables Information Accessibility
At the heart of MESA is a challenge that becomes more complicated as the volume and variety of scientific data continue to grow.
“Individual research teams often adopt idiosyncratic organizations for the data they collect,” Maccabe says. “MESA attacks this challenge using emerging agentic AI technologies to make the information easily accessible, independent of how the data is organized.”
MESA will use AI agents to examine scientific datasets, automatically generate and standardize metadata—the information that describes what data contains and how it is organized—and connect datasets with related information from other fields. Researchers will be able to ask questions and search across raw data, images, videos and other scientific resources rather than having to locate and reconcile them manually.
The goal is straightforward but ambitious: help researchers spend less time finding and preparing data and more time using it to make discoveries.
A Project Built Across Disciplines
At the U of A, the project itself reflects the kind of interdisciplinary science MESA is designed to enable.
The team brings together researchers and technical experts from the College of Information Science, College of Engineering and College of Science, along with the BIO5 Institute and the Office of Research and Partnerships.
In addition to Maccabe, the College of Information Science is represented by Assistant Professor of Practice Greg Chism, who will lead MESA’s education, outreach, training and documentation efforts.
The College of Engineering team includes David Ebert, U of A’s chief AI and data science officer and Computer Science Engineering Endowed Innovation Chair, and Jingdi Chen, assistant professor, whose work will contribute communications, 5G and 6G networking, and data security use cases.
From the College of Science, Assistant Professor Lei Cao will oversee database system design while Chi-Kwan “CK” Chan, associate astronomer and research professor at Steward Observatory, will provide expertise in astronomical data management through his work with the Event Horizon Telescope Collaboration.
Other U of A researchers, software engineers, data scientists and data curators will contribute expertise in cloud infrastructure, data management, research software development and human-computer interaction.
The five integrated layers of the MESA platform.
Image courtesy Barney Maccabe.
From Data Silos to Shared Infrastructure
For Chism, bringing such a wide range of disciplines together is fundamental to what MESA is designed to accomplish.
“MESA’s potential impact extends well beyond any single scientific field,” he says. “The project brings together use cases spanning astronomy, biology, environmental science, agriculture, computer science, engineering, communications, cybersecurity and public health. The commonality is that all disciplines differ substantially in their data, methods and research questions, but increasingly face the same fundamental challenge: how to find, integrate and responsibly use enormous amounts of distributed scientific data. In this, the interdisciplinarity of MESA is its truest contribution.”
Rather than creating another tool designed for one research field, MESA is intended to provide a common foundation that different scientific communities can use and adapt.
Its AI agents will help generate and harmonize metadata, connect related datasets and assist researchers in discovering information through technologies including retrieval-augmented generation, knowledge graphs and multi-step AI reasoning. Human-AI interaction, information visualization and user experience design will help make those capabilities accessible to researchers rather than requiring them to become experts in the underlying computing infrastructure.
The platform will be built on national research cyberinfrastructure, including CyVerse, the U of A-led platform originally created to support data-intensive plant science and now used by researchers across disciplines.
That shared infrastructure means advances developed for one type of scientific data could be applied elsewhere.
“MESA is designed so that advances made for one research community can transfer to many others,” Chism says. “It accomplishes this by developing shared, open-source infrastructure for AI-assisted data discovery, harmonization and analysis.”
Methods for organizing genomic data, for example, could inform environmental or chemical sciences. Techniques developed to work with massive astronomy datasets could be useful for agriculture or public health. Environmental data tools could help researchers working in areas ranging from civil engineering to epidemiology.
The two-year prototype will put that concept to the test across applications including astronomy simulations, life-science repositories, genomics and related biological research, precision agriculture, sensor networks and environmental data synthesis. MESA will work with established research communities and NSF-supported centers in environmental data science, molecular and cellular science and AI in agriculture, as well as the international Event Horizon Telescope Collaboration.
Preparing Researchers for AI-Enabled Science
Building the technology is only part of the challenge. Researchers also need to understand how to use AI-enabled tools effectively, reproducibly and responsibly.
That is where another key component of the College of Information Science’s work comes in.
Chism will lead MESA’s training and workforce-development efforts, including an Educator Fellows program and onboarding for early-adopter research communities. Training will address reproducible AI and data science, agentic AI, cloud-native open science and FAIR data practices designed to make scientific information more findable, accessible, interoperable and reusable.
The work reflects a central information science challenge: developing powerful new technologies is not enough. Researchers also need systems for organizing information, interfaces for interacting with it and practices that allow people and machines to interpret, share and reuse data across disciplinary boundaries.
For the researchers behind MESA, success would mean not simply making individual searches faster, but changing how scientists across fields work with one another and with the vast amounts of information that science now produces.
“The broader implication is a shift away from discipline-specific data silos toward a common foundation for AI-enabled, interdisciplinary science that allows researchers to spend less time preparing and locating data and more time asking questions that cross traditional disciplinary boundaries,” Chism says.
Learn more about the MESA project, or explore additional interdisciplinary research at the College of Information Science.