AITE Overview#
Description#
The Technology Test and Evaluation Division at the National Institute of Standards and Technology (NIST) has launched a new program to provide researchers with a sequestered testbed environment for the evaluation of AI model performance in a variety of meaningful tasks across diverse datasets, modalities, and domains. The Artificial Intelligence Technology Evaluation (AITE), currently in it’s initial phase, provides volunteer testing of AI models on blind data and its sequestered environment mitigates the risk of train/test data contamination to ensure rigorous, objective assessment. The infrastructure provided by NIST will provide common data, metrics and scoring to help developers understand the performance of their models.
The purposes of the effort include:
▪ serve as a neutral 3rd party for hosting such tests;
▪ be able to utilize data that is not publicly released, in order to remove the potential for train/test data contamination and enable the use of data where public distribution is undesired,
▪ measure the state of the art.
AITE will rely on engagement from participants in two different tracks, each offering distinct advantages:
▪ Data providers submit an original dataset in their domain that is inaccessible to others and a meaningful task to be performed on that dataset. Data providers will receive careful measurements of top models on their data conducting their task.
▪ Model providers submit AI models to be tested on the datasets and tasks. Model providers will learn how their models perform on an increasing number of datasets and tasks, and how their models perform relative to others on the same data using the same metrics, improving comparability while ensuring the evaluation data is not used for training any model.
Participation is open to all who wish to engage in one of the ways described above, and who can abide by the AITE Participation Agreement and rules. To request participation or ask questions, contact us at aite-poc@list.nist.gov. Results for all submitted systems are posted on the AITE website along with the identification of the submitting organizations. NIST summary reports containing general analysis are updated no less than once a year.
Currently, AITE consists of tasks for three use cases (1) Quantum Dot Control; (2) Human Genome Variant Curation; (3) Public Safety Visual Event Recognition.
Over time, AITE will build out multiple tasks under various themes (Quantum, Video, Natural Language Processing, …). Each task will have normal evaluation specification documents and submission protocols, and submissions will be scored and posted.
Available Tests#
Tests of various tasks, using various datasets, in various domains, and including various modalities.
Quantum dots are nanoscale semiconductor structures that can confine and control individual electrons, forming artificial charge islands with applications in single-electron electronics, quantum information processing, and sensing. Despite their promise, operating quantum dot devices remains a major experimental challenge. Device tuning has traditionally relied on expert intuition, trial-and-error procedures, and time-intensive measurements, making it as much an experimental “art” as a systematic workflow. These heuristic approaches become increasingly impractical as quantum dot systems grow in size and complexity. Scalable quantum technologies will require automated methods that can reliably characterize, tune, and control devices with reduced human intervention. The task described in this test captures a foundational capability for quantum dot control. As models improve, the test will expand to include increasingly complex tasks, providing a pathway for evaluating progress toward more autonomous quantum dot operation.
Dataset |
Test Plan |
Year |
Modalities: Input / Output |
# Trials |
Metric |
|---|---|---|---|---|---|
QDC Patches v1.0 |
2026 |
text and image / text |
641 |
Mean Squared Error |
The predominant approach to identifying biologically and clinically relevant genomic variation in humans is to align reads and call variants against a reference human genome. Accurate, high-resolution benchmark datasets have had fundamental impact in facilitating the development of sequencing technologies and bioinformatics tools as well as the validation of clinical tests. The creation of such benchmark datasets requires careful curation of genomic variants. The task described in this test represents a basic ability for genome variant curation and, as models demonstrate their abilities and improve over time, additional, increasingly complex tasks in support of human genome variant curation will be added to the test.
Dataset |
Test Plan |
Year |
Modalities: Input / Output |
# Trials |
Metric |
|---|---|---|---|---|---|
Genome Variant Visualization v1.0 |
2026 |
text and image / text |
10,000 |
Average Error Rate |
When public safety events occur, a swift and effective response is paramount, and failure to respond quickly and appropriately can result in major disruptions, as well as loss of lives and property that could otherwise be avoided. Real-time public safety imagery enables emergency personnel to optimally assess and respond to public safety events. However, the continuous monitoring, especially a large number, of imagery sources is a known challenge for humans. The task described in this test represents a basic ability to identify public safety events from imagery data. As models demonstrate their abilities and improve over time, increasingly complex tasks in support of public safety event recognition, potentially including additional modalities, will be added to the test.
Dataset |
Test Plan |
Year |
Modalities: Input / Output |
# Trials |
Metric |
|---|---|---|---|---|---|
Gumby V1.0 |
2026 |
text and image / text |
3,000 |
Detection Cost Function |
Models#
The following table lists the example model(s).
Model |
Date of Release |
Open Source(y/n) |
|---|---|---|
Fall, 2025 |
Y |