AI News Feed
Market watch
Companies

NVIDIA Releases Kumo Tabular, an Open Foundation Model for Tabular Prediction

NVIDIA has released Kumo Tabular, an open foundation model for tabular classification and regression that predicts new rows in a single forward pass without training or feature engineering. It is available on Hugging Face under a commercial-use license and ranks first on four benchmarks, according to NVIDIA.

Kumo Tabular is part of the NVIDIA Kumo Structured model collection. It comes in three sizes, from 28 million to 215 million parameters. It runs through NVIDIA's open-source library and is released under the OpenMDW-1.1 license for commercial use. Model code is available on GitHub, and model weights are available on Hugging Face.

The company positions the release against two decades of gradient-boosted trees, which remain common for enterprise machine learning on tabular data. Customer records, transactions, sensor logs, claims, and orders all live in tables, and predicting churn, default, demand, or price from them is among the most common machine learning tasks in industry. But the lifecycle around those models still involves collecting labels, engineering features, searching hyperparameters, validating, and deploying a model that learns each task from scratch.

Kumo Tabular instead uses in-context learning, an approach associated with large language models. Given a few examples in a prompt, a pretrained model can solve a task without updating a single weight. NVIDIA says the same idea applies to tables: a model pretrained on millions of tables can read a labeled table as its context and predict the labels of new rows directly.

The model is a Transformer built around the structure of a table, using column, row, and in-context attention as introduced in TabICL and TabPFN. To predict a label, it has to understand what each value means within its column, understand how the columns of a row interact, and relate context rows with existing labels to query rows with unknown labels.

In the first stage, cell embedding groups cells into tokens. Numerical and categorical values pass through Fourier features, sines and cosines of learned frequencies, with separate weights for each type. Missing values need no imputation and are treated specially. Every token in the context receives a label embedding.

Row embedding then turns each row into an embedding by alternating two kinds of attention multiple times. Column attention looks down a single column and learns what a value means in the distribution of its column, such as whether a 42 is typical or extreme, through induced self-attention; its cost grows linearly with the number of rows. Row attention looks across the tokens of a single row and learns how features interact, with rotary positions to tell columns apart. Four learnable CLS tokens join each row and act as the final readout. After this row compression, the cost of the final stage no longer depends on the number of columns.

A final Transformer operates on the row embeddings. Context rows attend to each other, while query rows attend to context rows only. Each prediction therefore depends only on the context and on the row itself, not on which other rows are scored alongside it. Because the context never looks at the queries, its keys and values are computed once and can be reused for follow-up predictions. Query rows use Test-GQA, which shrinks the cache that every prediction reads. A head turns each query row into class probabilities for classification and 999 quantiles for regression, from which a point prediction and an uncertainty estimate follow.

NVIDIA also describes length-aware attention temperature. Softmax attention spreads out as the number of keys grows, so attention that is sharp over a few hundred rows can dissolve over tens of thousands. Kumo Tabular scales every query by a temperature that grows with the logarithm of the number of keys, with a coefficient learned separately for each attention head. The result, according to NVIDIA, is attention that stays sharp as tables grow longer or wider.

Kumo Tabular was pretrained entirely on artificial tables. Each training table is sampled from a Structural Causal Model, or SCM. NVIDIA says the process first draws a configuration for the whole table, from its size and task to its mechanisms and missingness, and then uses a random causal graph to link hidden variables.