UH Mānoa Teams Join National AI Science Initiative
- Sophie Larsen

- Jul 31
- 13 min read
UH Mānoa reached Google News after two research teams entered a national AI program, but selection only starts a demanding nine-month test. One team will protect AI systems controlling critical infrastructure. The other will build a shared model for experiments searching for a rare nuclear process.
The projects join the U.S. Department of Energy’s inaugural Genesis Mission portfolio. Each qualifies for a Phase I award ranging from $500,000 to $750,000. Successful teams can later compete for Phase II awards reaching $15 million.
That progression creates the central tension. The program promises to turn federal data, national laboratory computers, and university expertise into faster scientific discovery. Yet both UH Mānoa projects operate where a plausible AI result is not enough. Grid operators need dependable responses under time pressure, while physicists need results that remain valid across different detectors.
Selection therefore does not establish scientific success. It gives the researchers nine months to show that their methods can survive operational constraints, unfamiliar data, and independent scrutiny.
The stakes extend beyond two laboratories in Hawaiʻi. Genesis involves 278 research efforts, 168 universities, and more than 340 institutions. The initiative is trying to replace isolated AI experiments with shared research infrastructure.
That ambition puts pressure on universities, national laboratories, and technology partners alike. They must prove that scale improves science without weakening security, traceability, or confidence in the results.
Two UH Mānoa Projects Enter the Genesis Mission
UH Mānoa is testing AI at two difficult boundaries: live infrastructure security and evidence-intensive particle physics.
The university announced the selections on July 27, following the Department of Energy’s July 22 portfolio announcement. Its project selection identifies Liuwan Zhu and Zepeng Li as the principal UH Mānoa researchers.
Zhu, an assistant professor of electrical and computer engineering, leads STRATOS. The name stands for Security and Trust Runtime Architecture for Time-critical Operational Science.
STRATOS will focus on AI systems used in real-time operations. Its target environments include electrical grid management and forecasting, where a late or incorrect response can have physical consequences.
The project is not simply another cybersecurity monitoring tool. It aims to observe an AI system while that system is making time-sensitive operational decisions. The platform must then detect suspicious behavior and support a response without disrupting legitimate work.
The research team plans to test STRATOS across two sites. One is the UH Mānoa campus microgrid. The other uses real-time computing systems at Argonne National Laboratory.
A microgrid is a localized energy network that can manage electricity generation, storage, and consumption. It provides a concrete environment for studying how an attack or model failure might affect operational decisions.
Old Dominion University and Argonne National Laboratory are collaborating on the project. That structure gives the team access to expertise beyond campus, but it also raises the standard for interoperability.
Li, a professor of physics and astronomy, leads the second project. His team will develop an AI foundation model for searches involving neutrinoless double beta decay.
A foundation model is trained on broad data so it can support multiple related tasks. In this case, the model is intended to work across experiments that use different detector technologies and analysis methods.
Neutrinoless double beta decay is a hypothetical nuclear process. Observing it would indicate that neutrinos are their own antiparticles, changing scientists’ understanding of matter and fundamental physics.
The process, if it occurs, is extremely rare. Researchers must separate a possible signal from background events recorded by sensitive detectors.
Different experiments approach that challenge with specialized hardware, software, and data conventions. Li’s project seeks a shared AI framework that learns across those differences.
The two projects appear unrelated at first. One concerns cyber threats around an electrical grid, while the other studies a question about the structure of the universe.
They share a harder problem beneath the surface. Both must establish when an AI system deserves trust in an environment where mistakes carry more weight than a poor benchmark score.
For STRATOS, trust means detecting threats without creating unsafe delays or false alarms. For the physics project, it means preserving the distinctions between experiments while finding patterns across them.
The Google News headline captures the selection. It does not capture how much validation remains before either approach becomes reusable scientific infrastructure.
Why AI-Driven Science Is Moving Beyond Isolated Models
Genesis treats AI-driven science as an infrastructure problem, not a contest to produce the largest general-purpose model.
President Donald Trump created the Genesis Mission through Executive Order 14363 on November 24, 2025. The order directs the Energy Department to establish an integrated American Science and Security Platform.
The executive order describes a system connecting high-performance computing, scientific datasets, AI models, analytical tools, and experimental facilities. It also calls for AI agents that can evaluate outcomes and automate research workflows.
This design reflects a practical limitation in scientific AI. A model alone cannot accelerate an experiment if researchers cannot access suitable data, computing capacity, instruments, or verification procedures.
Scientific data also differs from the public text and images commonly used to train commercial models. It can include sensor streams, simulation outputs, detector events, proprietary measurements, and classified information.
Those sources come with distinct access controls and uncertainty. A useful system must record where data originated, how it changed, and which assumptions shaped the analysis.
Genesis attempts to coordinate those components through national laboratories and outside partners. The Department of Energy operates laboratories with specialized computers, instruments, and decades of scientific data.
The initiative also reaches beyond the department. More than 15 federal agencies announced contributions to Genesis in July 2026. Their involvement spans funding, datasets, facilities, and research programs.
The National Science Foundation separately invited proposals supporting AI-enabled scientific discovery. Its research invitation highlights autonomous laboratories, scientific datasets, physical constraints, and AI-assisted hypothesis generation.
That language matters because it shifts the goal from prediction toward participation in the research cycle. The system is expected to help propose, analyze, and refine experiments.
However, greater participation creates greater verification demands. A model that recommends the next experiment influences which evidence researchers collect. An error can therefore propagate beyond a single output.
The funding structure recognizes that uncertainty. Under the Department of Energy’s funding program, Phase I projects receive nine months and $500,000 to $750,000. Phase II awards range from $6 million to $15 million over three years.
Phase I is therefore a feasibility test, not a smaller version of full deployment. Teams need to show that their method addresses a genuine scientific or operational bottleneck.
They must also demonstrate that the approach can move beyond a single curated dataset. A system that only works in one laboratory offers limited value to a national platform.
For UH Mānoa, the national structure creates opportunity and pressure. The teams gain access to collaborators and infrastructure that would be difficult to assemble locally.
In return, their work must support a broader mission. STRATOS needs lessons that apply beyond the campus microgrid. The physics model needs to transfer across detector technologies without erasing meaningful differences.
This is why the UH selections matter beyond Google News visibility. They place a geographically remote public university inside an effort to define how shared AI systems will support American science.
The selection also tests whether a national platform can incorporate specialized researchers without concentrating every important capability at a few large institutions.
Google News Attention Hides a Security Versus Scale Conflict
The main contest is between shared AI infrastructure and the safeguards required by high-consequence science.
Genesis depends on scale. Its central logic connects more data, more computing resources, more models, and more institutions.
Security often pushes in the opposite direction. Sensitive systems rely on restricted access, isolated environments, narrow permissions, and carefully controlled changes.
STRATOS sits directly inside this conflict. A security platform needs enough visibility to recognize abnormal behavior across an AI workflow. Granting that visibility can itself create privileged access that must be protected.
The problem becomes harder in real-time operations. Conventional security processes can pause a system, quarantine data, or request human review. An electrical control environment may not tolerate the same delay.
A threat detector can also harm operations through false positives. If it repeatedly labels legitimate behavior as hostile, operators may disable it or ignore its warnings.
False negatives carry the opposite risk. A system might continue trusting manipulated data or malicious model behavior because the attack resembles ordinary variation.
STRATOS must therefore balance detection, response speed, and operational continuity. Improving one measure does not automatically improve the others.
Its dual-site test is important for that reason. A campus microgrid offers a physical operating environment, while Argonne contributes national laboratory computing and research expertise.
Testing across both locations can expose assumptions hidden by a single deployment. Network timing, data formats, hardware, and access policies can all affect performance.
Yet two sites cannot represent every energy system. Utilities vary in size, equipment age, regulatory requirements, communications technology, and tolerance for automated intervention.
A successful demonstration would show that the architecture functions under defined conditions. It would not prove universal readiness for critical infrastructure.
The foundation model for neutrinoless double beta decay faces a related scale problem. Shared learning can help experiments benefit from one another’s data. However, detector differences are not merely inconvenient formatting issues.
Each detector technology captures signals through its own physical mechanism. Those choices affect resolution, background noise, calibration, and systematic uncertainty.
A model that combines data too aggressively can learn shortcuts tied to a particular instrument. It might then perform well during internal testing but fail when applied elsewhere.
A shared framework needs mechanisms to represent those differences explicitly. It also needs evaluation procedures that reveal when the model is uncertain or operating outside its training distribution.
The issue is especially serious in rare-event physics. Researchers may observe only a tiny number of candidate events among a much larger background.
An AI system can help classify events and identify patterns. It cannot replace the statistical and experimental standards required to claim a new physical process.
That distinction separates useful assistance from automated authority. AI can help researchers find where to look, but the evidence must remain open to conventional scientific review.
The broader Genesis program faces the same tension. Its national platform needs common interfaces and reusable models. Scientific credibility depends on preserving domain-specific context.
The Energy Department says consortium members have committed more than $800 million in support. The partner commitments include computing resources, cloud infrastructure, model access, scientific expertise, and direct funding.
Those contributions expand what researchers can attempt. They also introduce questions about access, intellectual property, vendor dependence, and responsibility when a shared component fails.
Scale is therefore not an unconditional advantage. Every new dataset, model, institution, and computing service adds capability alongside another trust boundary.
Google News can make the initiative appear like a single coordinated national project. In practice, its credibility will emerge from many narrower decisions about permissions, validation, documentation, and human review.
Shared Models Must Preserve Scientific Differences
A common AI framework becomes valuable only when it transfers knowledge without flattening the experiments that produced it.
Li’s project targets a long-standing coordination challenge in rare-event physics. Multiple research collaborations pursue the same underlying phenomenon using different detector materials and analysis pipelines.
Each experiment has reasons for its design choices. Those differences provide independent ways to test a claim and expose possible measurement errors.
A shared foundation model can create efficiency by learning general features across experiments. It might help identify event shapes, improve background rejection, or transfer representations to a smaller dataset.
However, transfer is not automatically valid. A pattern learned from one detector might reflect its geometry, calibration process, or simulation assumptions.
The model must distinguish general physics from instrument-specific artifacts. Otherwise, apparent collaboration can produce correlated errors across projects.
This challenge has no simple benchmark. Conventional machine-learning evaluation often separates one dataset into training and test samples drawn from similar conditions.
A scientific model needs stronger tests. Researchers should evaluate it on different detector configurations, changed background conditions, and data withheld by institution or time period.
They also need comparison points outside AI. Established analysis pipelines provide a baseline for measuring whether the shared model adds sensitivity without introducing unacceptable uncertainty.
Interpretability matters here, but it should not become a vague demand for a simple explanation. The relevant question is whether researchers can trace a result to data, model behavior, and known physical assumptions.
Reproducibility creates another requirement. Independent teams need enough information to run the analysis, inspect preprocessing choices, and understand model updates.
A model can change as teams add new data or revise calibration. Version control and provenance become part of the scientific record, not merely software administration.
These issues resemble challenges faced by organizations building an AI knowledge base. Shared information gains value only when users can locate its origin, context, and latest revision.
Scientific collaborations face stricter consequences, but the information problem is recognizable. A generated answer or classification is less useful when nobody can determine which evidence shaped it.
The physics project also tests whether foundation-model economics translate into research. Commercial AI often benefits from broad training followed by adaptation to individual tasks.
Scientific experiments may have fewer labeled examples, tighter uncertainty requirements, and data that cannot move freely between institutions. Training costs also matter when models require repeated evaluation.
A useful framework must produce enough scientific benefit to justify that complexity. Faster analysis alone is not sufficient if researchers spend comparable time validating opaque behavior.
The project’s nine-month Phase I period limits what can be established. The team can develop the architecture, test transfer between selected detector contexts, and identify failure modes.
It cannot settle the neutrino question within that window. It also cannot prove that one model architecture will serve every present and future experiment.
The realistic milestone is narrower. The team needs evidence that shared learning improves a meaningful analysis task while preserving detector-specific uncertainty.
That would support a Phase II case. It would also offer a template for other scientific fields where institutions collect related but nonidentical data.
Failure would still produce useful information if documented carefully. The team might find that data harmonization, simulation differences, or governance creates a larger barrier than model design.
Such a result would challenge a common assumption behind national AI programs. More connected infrastructure does not automatically create compatible evidence.
The same lesson applies to STRATOS. A common runtime security architecture must adapt to different systems without treating every difference as a threat.
Both projects therefore test whether shared AI can respect local context. Their domains differ, but their verification problem is closely aligned.
What the First Nine Months Cannot Prove
Phase I can establish technical feasibility, but it cannot validate nationwide deployment or a fundamental discovery.
The funding announcement creates an easy risk of overstatement. Awards up to $750,000 sound substantial, while possible Phase II support reaches $15 million.
Those figures describe eligibility and competition stages. They do not mean that every selected project has received the maximum amount or will advance.
The projects must meet milestones during a compressed development period. Nine months is enough to build and test a focused prototype, but it is short for evaluating rare failures.
Cybersecurity systems often appear dependable until attackers change tactics. Scientific models can also perform well on familiar data while failing after an instrument or dataset shifts.
STRATOS needs adversarial testing, meaning deliberate attempts to confuse or evade the platform. The relevant tests should include manipulated inputs, compromised components, timing pressure, and ordinary operational anomalies.
A security demonstration should report more than detection accuracy. Response latency, false-alarm rates, service interruptions, and recovery behavior will shape practical value.
Human factors deserve equal attention. Operators need clear information about why the system raised an alert and what action it recommends.
An architecture that produces frequent, unexplained warnings can increase workload. In a time-critical setting, that burden can become a safety problem.
The dual-site platform will reveal some integration issues. It cannot reproduce the full diversity of power infrastructure or the incentives of real attackers.
The team should therefore frame any result around its test conditions. Claims about protecting critical infrastructure should remain narrower than claims about protecting a particular evaluated environment.
Li’s foundation model requires similar caution. Better classification performance on selected datasets would not constitute evidence of neutrinoless double beta decay.
The model could instead help experiments use their data more efficiently. Any candidate signal would still require extensive statistical review and independent confirmation.
Researchers must also guard against simulation bias. Rare-event experiments rely heavily on simulated examples because confirmed examples of the target process do not exist.
If simulations omit an important background process, an AI model can learn a distinction that disappears in real detector data.
Cross-experiment training might reduce some blind spots by exposing the model to varied conditions. It might also spread a shared modeling assumption across several analyses.
That tradeoff makes independent baselines essential. A common model should complement, rather than eliminate, distinct analysis methods during validation.
Governance remains another uncertainty. The public announcements describe collaboration and infrastructure, but they provide limited detail about long-term model ownership and data access.
Questions include who approves model updates, who investigates errors, and which results can be shared outside participating institutions.
National-security applications add further restrictions. Some datasets or system details cannot be fully open, even when transparency would help scientific review.
Genesis must find workable boundaries between openness and controlled access. Those boundaries will differ by project, weakening the idea of one uniform platform.
Funding durability also matters. Phase I teams can build prototypes, while durable infrastructure requires maintenance after a research award ends.
Models need monitoring, security patches, updated documentation, and support for new data. A neglected shared model can become less reliable even if its original paper remains influential.
These concerns do not make the projects poor investments. They define what credible progress should look like.
The useful standard is not whether a team declares success. It is whether outside researchers can understand the test, reproduce relevant results, and identify where the system still fails.
That distinction should guide readers who encounter the story through Google News. Selection is evidence that reviewers found the proposals promising. It is not independent confirmation of their eventual claims.
Three Signals Will Show Whether the Bet Is Working
The next evidence should come from tests, transfer results, and Phase II decisions, in that order.
The first signal is a detailed STRATOS evaluation across the UH Mānoa microgrid and Argonne systems. The most informative results will describe attacks tested, response time, false alarms, and operational disruption.
A simple claim that the system detected threats would offer little basis for judgment. Readers need to know whether it performed under realistic timing and whether operators could act on its warnings.
Evidence that STRATOS works across both sites would strengthen the case for a reusable security architecture. Large performance differences between the sites would suggest that local system design remains the dominant factor.
The second signal is cross-detector transfer from Li’s foundation model. The team should show whether a model trained with multiple detector contexts improves a defined task on data it did not encounter during training.
The result should include uncertainty and comparisons with established pipelines. Performance gains without those controls would weaken confidence.
A successful transfer test would support the broader premise behind AI-driven science. It would show that shared representations can preserve enough physical context to help separate experiments.
Failure to transfer would not invalidate AI in physics. It would indicate that detector differences require more specialized models or better data harmonization.
The third signal is the Phase II selection process. Awards reaching $15 million would indicate which Phase I projects convinced reviewers that their methods deserve broader deployment.
Advancement alone will not prove effectiveness. The selection criteria, proposed milestones, and partner commitments will reveal what the program values after its first feasibility cycle.
If projects advance mainly because they promise larger models or more infrastructure, verification may remain secondary. If reviewers demand cross-site tests, reproducible evidence, and documented failure modes, Genesis will gain credibility.
The national program should also reveal how its platform connects the 278 research efforts. Shared standards for data provenance, model evaluation, and security would matter more than a central project directory.
The Department of Energy and its partners have assembled money, computing resources, institutions, and political attention. Coordinating those assets is now the difficult part.
UH Mānoa has a meaningful position in that test. Its projects address two situations where AI cannot rely on persuasive output alone.
STRATOS must make decisions useful under operational pressure. The physics foundation model must help analyze evidence without distorting its experimental origins.
Readers should resist treating the selections as either proof of an AI research transformation or another empty funding announcement. The first results can establish something more practical.
They can show whether national-scale AI collaboration improves narrowly defined scientific work while preserving accountability. That would be a significant achievement, even before a major discovery or infrastructure deployment.
The next Google News headline should therefore matter less than the evidence behind it. Watch for measurable security performance, cross-detector validation, and transparent Phase II milestones.
Researchers, developers, and enterprise AI buyers can apply the same standard to their own systems. Ask where the data came from, how the model behaves outside familiar conditions, and who reviews consequential errors.
Those questions turn AI adoption into an evidence problem rather than a branding exercise. UH Mānoa’s nine-month tests now have an opportunity to supply some of the answers.


