AI News Feed
Market watch
Companies

NVIDIA Releases Kumo Tabular, Open Tabular Foundation Models With Commercial-Use Weights

NVIDIA released Kumo Tabular, open tabular foundation models for one-pass prediction with commercial-use weights.

Kumo Tabular comes in Small, Medium, and Large versions, spanning about 28 million to 215 million parameters. The SDM code is Apache-2.0 and requires Python 3.11+ and PyTorch 2.7+, with examples targeting a CUDA GPU. SDM is a GPU-native library for structured-data foundation models and preprocessing. Besides Kumo Tabular, it includes TabICLv2, Google’s TabFM, and KumoRelational for multi-table data. The models share one in-context learning interface built on a TableTensor container; the library also handles preprocessing, ensembling, and many-class prediction.

Kumo Tabular is a Transformer built around the structure of a table. It uses column, row, and in-context attention, as introduced in TabICL and TabPFN. The pipeline has three stages. In cell embedding, numerical and categorical values pass through learned Fourier features with separate weights per type, and missing values require no imputation. In row embedding, column attention uses induced self-attention so cost grows linearly with rows; row attention with rotary positions learns feature interactions, and four learnable [CLS] tokens compress each row. In in-context learning, a final Transformer runs over row embeddings. Context rows attend to each other, while query rows attend only to context rows. Because the context never sees the queries, its keys and values are computed once and reused. The head outputs class probabilities, or 999 quantiles for regression, giving a point prediction plus an uncertainty estimate. Kumo Tabular also scales each query by a temperature that grows with the log of the key count, with the coefficient learned per attention head.

The models are pretrained entirely on synthetic tables sampled from structural causal models. A random causal graph links hidden variables through linear maps, small neural networks, trees, or Gaussian processes. The generator injects messy, real-world patterns such as missing values, high-cardinality categories, heavy-tailed targets, and conflicting duplicate rows. Training ran in three stages, similar to TabICLv2. Context grew from 1,024 rows to 60,000 rows, with up to 100 columns. Small, Medium, and Large saw about 35 million, 71 million, and 137 million artificial tables. Classification and regression are trained as separate models. NVIDIA says the training recipe and data generators will be released soon.

With default settings, Kumo Tabular ranks first overall on TabArena with an Elo of 1950, according to NVIDIA. The NVIDIA team reports it runs 17 times faster than LimiX-2 on a single RTX 6000 Pro. All three sizes sit on the accuracy and inference-time Pareto front. On BeyondArena, it places first with an Elo of 1418 and an Improvability score of 7.78 percent. On TALENT, it has the top overall ranking, with average ranks of 6.67 for accuracy, 3.98 for log-loss, and 4.22 for RMSE. On ScoringBench, Large and Medium rank first and second on average rank.

Licensing separates Kumo Tabular from several competitors, MarkTechPost reported. Kumo Tabular has about 28 million to 215 million parameters across three sizes, while TabICLv2 has 27.55 million parameters for classification and 28.54 million for regression, LimiX-2 has 400 million, and TabFM has about 1.64 billion. TabPFN-3 does not list parameter counts in its documentation. All five handle classification and regression, though LimiX-2 also handles imputation. Kumo Tabular and TabICLv2 support 10 native classes per pass, with ECOC or hierarchical methods for more; TabPFN-3 supports 160, TabFM has a hard limit of 10, and LimiX-2 is not specified.

On weights licensing, Kumo Tabular uses OpenMDW-1.1, TabICLv2 uses BSD-3-Clause, TabPFN-3 uses the TABPFN-3 License v1.0, LimiX-2 uses the StableAI LimiX Non-Commercial license, and TabFM uses the TabFM Non-Commercial v1.0 license. Commercial use of weights is allowed for Kumo Tabular and TabICLv2, requires a paid license for TabPFN-3, and is not allowed for LimiX-2 and TabFM. Kumo Tabular, TabICLv2, and TabFM run in NVIDIA SDM; TabPFN-3 and LimiX-2 do not. MarkTechPost called the license row the real differentiator, noting that TabPFN-3, LimiX-2, and TabFM weights carry non-commercial terms, while Kumo Tabular and TabICLv2 are the permissive options and Kumo Tabular leads the benchmarks NVIDIA reports.

To get started, users can install the library and pass a DataFrame through TableTensor, following the model card. The material lists pip install structured-data-models and an example using scikit-learn’s breast cancer dataset, sdm.TableTensor.from_pandas, and sdm.models.KumoTabular.