{"article":{"slug":"tabpfn-vs-xgboost-benchmark-measured-on-an-rtx-4070-ti","title":"TabPFN vs XGBoost: benchmark measured on an RTX 4070 Ti","subtitle":null,"summary":"The claim behind TabPFN and TabICL is that they predict on a table without ever training on it and still beat tuned boosting. I measured it on fourteen datasets from the Grinsztajn benchmark, with the same split and the same clock for everyone. The one that does not train wins, the advantage holds up to 32,000 rows instead of breaking, and the most-cited model can no longer be downloaded without a","content_type":"blog_post","language":"en","canonical_url":"https://efraingaray.com/en/blog/tabpfn-vs-xgboost/","author":{"name":"Efraín Garay","url":"https://efraingaray.com/","person_slug":null,"person_url":null},"authored_by":"human","publisher":{"name":"Efraín Garay","url":"https://efraingaray.com/","listing_slug":null,"listing":null},"topics":[{"name":"Machine Learning","slug":"machine-learning","url":"https://listedarticles.com/topics/machine-learning"},{"name":"Benchmarks","slug":"benchmarks","url":"https://listedarticles.com/topics/benchmarks"},{"name":"AI","slug":"ai","url":"https://listedarticles.com/topics/ai"},{"name":"Research","slug":"research","url":"https://listedarticles.com/topics/research"}],"about_listings":[],"cover_image_url":null,"license":"all-rights-reserved","word_count":3822,"reading_minutes":17,"published_at":"2026-08-17T12:00:00.000Z","added_at":"2026-09-28T06:19:36.320Z","updated_at":"2026-09-28T06:19:36.320Z","added_via":"api","contributor":{"type":"agent","name":"ListedStartups Using Bot","registered":false},"profile_url":"https://listedarticles.com/articles/tabpfn-vs-xgboost-benchmark-measured-on-an-rtx-4070-ti","markdown_url":"https://listedarticles.com/articles/tabpfn-vs-xgboost-benchmark-measured-on-an-rtx-4070-ti.md","example":false,"citation":"Efraín Garay, Efraín Garay. \"TabPFN vs XGBoost: benchmark measured on an RTX 4070 Ti.\" 17 Aug 2026. https://efraingaray.com/en/blog/tabpfn-vs-xgboost/ (all-rights-reserved)","access":{"human_view":"preview","full_text_available":true,"source_url":"https://efraingaray.com/en/blog/tabpfn-vs-xgboost/"},"body_markdown":"# TabPFN and TabICL against tuned XGBoost: the model that does not train won on fourteen tables out of fourteen\n\nThe claim behind TabPFN and TabICL is that they predict on a table without ever training on it and still beat tuned boosting. I measured it on fourteen datasets from the Grinsztajn benchmark, with the same split and the same clock for everyone. The one that does not train wins, the advantage holds up to 32,000 rows instead of breaking, and the most-cited model can no longer be downloaded without an account.\n\nThe claim has been going around for months and it is concrete enough to be measurable: a tabular foundation model predicts on a table **without ever having trained on it** and still beats tuned boosting.\n\nIf that is true, half a decade of practice changes shape. Searching hyperparameters stops being a mandatory step and becomes a luxury that sometimes does not pay off.\n\nSo I put it to the test on my own card, with fourteen datasets, four contenders and the same stopwatch for everyone.\n\n## What exactly these models do\n\nA tabular foundation model is pretrained on millions of synthetic tables generated on purpose. When a new table arrives, it does not adjust a single weight: it receives the training rows **as context** and produces the predictions in one forward pass.\n\nIt is the same in-context learning idea we already know from language models, moved from words to columns. That is why the verb “train” sits oddly: the code still calls it `fit`, but inside there is no gradient descent, there is a copy of data to the card.\n\nThat also explains why the cost shows up where you do not expect it. Fitting is nearly free and **prediction** is what pays, exactly the reverse of a tree.\n\nPut that way it sounds abstract, so here is the full journey of one row: a request arrives with an empty cell, the model weighs it against everything that already happened, resolves it in one pass and returns a future. Pick any of the six cases I measured to see it with their real data:\n\nThe request\n\nThe context\n\nA single pass\n\nThe prediction\n\nA credit application arrives\n\n2 100 applications already settled · 10 columns`clf_num/credit.csv`\n\nTabICL · nothing tuned · **0.8 s**\n\nWill they repay?\n\nrepaysdefaults\n\nAUC **0.8528**\n\ncredit\n\nA contact enters the list\n\n2 100 calls already made · 7 columns`clf_num/bank-marketing.csv`\n\nTabICL · nothing tuned · **0.8 s**\n\nIs it worth calling them?\n\nsigns updoes not\n\nAUC **0.8699**\n\nbank-marketing\n\nA patient is discharged\n\n2 100 previous discharges · 7 columns`clf_num/Diabetes130US.csv`\n\nTabICL · nothing tuned · **0.8 s**\n\nWill they be readmitted?\n\nreturnsdoes not\n\nAUC **0.6484**\n\nDiabetes130US\n\nA case is assessed\n\n2 100 cases already closed · 11 columns`clf_cat/compas-two-years.csv`\n\nTabICL · nothing tuned · **0.6 s**\n\nWill they reoffend?\n\nreoffendsdoes not\n\nAUC **0.7329**\n\ncompas-two-years\n\nA market period closes\n\n2 100 previous periods · 7 columns`clf_num/electricity.csv`\n\nTabICL · nothing tuned · **0.8 s**\n\nDoes the price go up or down?\n\nupdown\n\nAUC **0.8873**\n\nelectricity\n\nA candidate molecule arrives\n\n2 100 molecules already assayed · 419 columns`clf_num/Bioresponse.csv`\n\nTabICL · nothing tuned · **6.0 s**\n\nDoes it trigger a biological response?\n\nactiveinert\n\nAUC **0.8667**\n\nBioresponse\n\nThe request\n\nThe context\n\nA single pass\n\nThe prediction\n\nA credit application arrives\n\n2 100 applications already settled · 10 columns`clf_num/credit.csv`\n\nTabICL · nothing tuned · **0.8 s**\n\nWill they repay?\n\nrepaysdefaults\n\nAUC **0.8528**\n\ncredit\n\nA contact enters the list\n\n2 100 calls already made · 7 columns`clf_num/bank-marketing.csv`\n\nTabICL · nothing tuned · **0.8 s**\n\nIs it worth calling them?\n\nsigns updoes not\n\nAUC **0.8699**\n\nbank-marketing\n\nA patient is discharged\n\n2 100 previous discharges · 7 columns`clf_num/Diabetes130US.csv`\n\nTabICL · nothing tuned · **0.8 s**\n\nWill they be readmitted?\n\nreturnsdoes not\n\nAUC **0.6484**\n\nDiabetes130US\n\nA case is assessed\n\n2 100 cases already closed · 11 columns`clf_cat/compas-two-years.csv`\n\nTabICL · nothing tuned · **0.6 s**\n\nWill they reoffend?\n\nreoffendsdoes not\n\nAUC **0.7329**\n\ncompas-two-years\n\nA market period closes\n\n2 100 previous periods · 7 columns`clf_num/electricity.csv`\n\nTabICL · nothing tuned · **0.8 s**\n\nDoes the price go up or down?\n\nupdown\n\nAUC **0.8873**\n\nelectricity\n\nA candidate molecule arrives\n\n2 100 molecules already assayed · 419 columns`clf_num/Bioresponse.csv`\n\nTabICL · nothing tuned · **6.0 s**\n\nDoes it trigger a biological response?\n\nactiveinert\n\nAUC **0.8667**\n\nBioresponse\n\n### Default risk\n\nthe foundation model wins\nBanking’s most repeated case: deciding who gets lent to.\n\n| dataset | size | TabICL | TabPFN | tuned XGB | \n|---|---|---|---|---|\n| credit | 3000 × 10 | 0.7667 | 0.7578 | 0.7533 | \n| heloc | 3000 × 22 | 0.7222 | 0.7300 | 0.7078 | \n| default-of-credit | 3000 × 20 | 0.6956 | 0.6967 | 0.6944 | \n\nAll three datasets go to the foundation model. On heloc TabPFN takes 0.022 and TabICL 0.014: in credit risk that is not decoration.\n\nContext: `clf_num/credit.csv` from the inria-soda/tabular-benchmark suite, trimmed to 3,000 rows with seed 0 and split 70/30 with stratification.\n\nsha256 `22a759296600c39a56884dafd84eb346be018c537ff096adc8c2ae7e0520f2a9`\n\n### Sales campaign\n\nthe foundation model wins\nWho to call first when there are a thousand contacts and time for a hundred.\n\n| dataset | size | TabICL | TabPFN | tuned XGB | \n|---|---|---|---|---|\n| bank-marketing | 3000 × 7 | 0.7944 | 0.7967 | 0.7833 | \n\nThe only case where TabPFN ends up ahead of TabICL. Both beat tuned boosting.\n\nContext: `clf_num/bank-marketing.csv` from the inria-soda/tabular-benchmark suite, trimmed to 3,000 rows with seed 0 and split 70/30 with stratification.\n\nsha256 `3433fecd416ce949692442882669881746e8a6b24b904a5f808cdcb196443e7e`\n\n### Clinical readmission\n\nthe foundation model wins\nA real hospital record: which patient gets readmitted.\n\n| dataset | size | TabICL | TabPFN | tuned XGB | \n|---|---|---|---|---|\n| Diabetes130US | 3000 × 7 | 0.5911 | 0.5933 | 0.5900 | \n\nOn accuracy they nearly tie, but the area under the curve opens up sharply: 0.6484 against 0.6264. The risk ordering, which is what a triage uses, improves considerably more than accuracy suggests.\n\nContext: `clf_num/Diabetes130US.csv` from the inria-soda/tabular-benchmark suite, trimmed to 3,000 rows with seed 0 and split 70/30 with stratification.\n\nsha256 `9be384ad7edbb9a98b509adb9cc08578e63bd87545b05d97d5147aa15b71385a`\n\n### Recidivism\n\nthe foundation model wins\nThe dataset that opened the debate on algorithmic bias in the courts.\n\n| dataset | size | TabICL | TabPFN | tuned XGB | \n|---|---|---|---|---|\n| compas-two-years | 3000 × 11 | 0.6733 | 0.6767 | 0.6700 | \n\nThe foundation model wins, and it is worth saying that a better model here does not make the use legitimate: this dataset’s argument was never about accuracy.\n\nContext: `clf_cat/compas-two-years.csv` from the inria-soda/tabular-benchmark suite, trimmed to 3,000 rows with seed 0 and split 70/30 with stratification.\n\nsha256 `7e0bcef09ae633ad81e26b0a8f8f96dbc4ac3c29b9da0760d38ef2a75adde8d8`\n\n### Electricity demand\n\na tie\nConsumption and price series from an electricity market.\n\n| dataset | size | TabICL | TabPFN | tuned XGB | \n|---|---|---|---|---|\n| electricity | 3000 × 7 | 0.8178 | 0.8111 | 0.8178 | \n\nThe tie. TabICL matches tuned boosting to the fourth decimal and TabPFN falls below. It is one of the two cases where the advantage does not show up.\n\nContext: `clf_num/electricity.csv` from the inria-soda/tabular-benchmark suite, trimmed to 3,000 rows with seed 0 and split 70/30 with stratification.\n\nsha256 `d6005007c4b1f7ba88cf91cb1f231969cce3e49c98d445338c0ab0342ead5d7e`\n\n### Molecular screening\n\ndepends on the model\n419 columns of chemical descriptors: the wide table.\n\n| dataset | size | TabICL | TabPFN | tuned XGB | \n|---|---|---|---|---|\n| Bioresponse | 3000 × 419 | 0.7922 | 0.7567 | 0.7744 | \n\nThis is where TabPFN breaks: 0.7567 is worse than even untuned XGBoost, and it costs 18.1 seconds. TabICL holds up and comes first. The table’s width, not its length, is what squeezes.\n\nContext: `clf_num/Bioresponse.csv` from the inria-soda/tabular-benchmark suite, trimmed to 3,000 rows with seed 0 and split 70/30 with stratification.\n\nsha256 `bd3d277821eb949df41219363f0386549019249776fccf7a9f08d3bec10b8727`\n\n## What this is for, concretely\n\nThe six cases in the diagram are not brochure examples: they are the datasets I measured with. But it is worth widening the map, because “tabular foundation model” sounds like a laboratory and the problem it solves is one of the most common there is.\n\nA table is a spreadsheet: rows that are cases and columns that are attributes. And the task is always the same, **predict one column from the others**:\n\n- A customer table with tenure, plan, usage and complaints, to estimate **which ones are going to leave next month** .\n- A transaction log with amount, merchant, hour and country, to flag **which ones are fraud** .\n- A history of credit applications with income, debt and payment behaviour, to estimate **who is not going to repay** . Three of the datasets I measured are exactly that:`credit` ,`heloc` and`default-of-credit` .\n- Sensor readings from a machine, to anticipate **when it is going to break** .\n- Patient records with symptoms and lab results, to prioritize **who gets seen first** .`Diabetes130US` , another of the datasets, is a real hospital record.\n\nThat is half the data work done in any company. It is also the ground where deep learning had been losing for years: for tables, a good gradient-boosted tree ensemble was still the right answer, and the research that assembled these datasets was titled exactly that way, asking why trees still won.\n\n### What changes if the promise holds\n\nToday, solving any of those cases has a ritual: prepare the data, choose a model, **search hyperparameters**, validate, repeat. The search is the boring part and the one that eats machine and human hours. In this benchmark, XGBoost’s search took up to 50 seconds per dataset; on a real problem with more rows and more combinations, it is minutes or hours.\n\nA tabular foundation model proposes skipping that ritual entirely. You hand it the table, it answers in a second, and that is that. No tree depth to choose, no learning rate, no cross-validation to decide among twenty-five candidates.\n\nPut in concrete terms: instead of spending the afternoon tuning a churn model, you have an answer in the time it takes to make coffee, and only then do you decide whether anything is worth refining. For exploring a table that just arrived, or for having an honest baseline before investing time, it is hard to beat.\n\nThe question, then, is not whether the idea is attractive. It is whether the result holds up when measured.\n\n## The benchmark’s rules\n\nA badly built benchmark says whatever you want. These are the rules I imposed on myself before seeing a single number:\n\n**The data is third-party and from the field where this argument is fought.** I used Grinsztajn’s tabular benchmark, the suite gathered for the work asking why trees still beat deep learning on tables. Choosing the datasets yourself is the easiest way to manufacture the result you like.\n\n**XGBoost goes in tuned, not for decoration.** The claim says “tuned boosting”, so comparing against factory parameters would be a straw man. I ran two versions: one with fixed, reasonable parameters and another with a random search over twenty-five combinations and three-fold cross-validation.\n\n**Same split, same seed, same columns for everyone.** Categoricals are integer-encoded. It is not the best possible treatment, but it is identical for all four, which is what makes the number comparable.\n\n**Five seeds per dataset, and the median is reported.** A single split cannot tell signal from luck.\n\n**Prediction time includes two inference passes**, the class one and the probability one, because the benchmark needs both to compute accuracy and area under the curve. That makes things more expensive precisely for the foundation models, which is where their cost lives, so the seconds I publish for them are inflated in nobody’s favour.\n\n**Two cores are left free.** The machine has other work on it, and a benchmark that eats the whole CPU measures the fight with the scheduler, not the model.\n\nEverything ran on a 16 GB RTX 4070 Ti SUPER, with fourteen cores for XGBoost.\n\n## Three stumbles before the first number\n\nPublishing only the final table would be lying by omission. This is what it cost to get there.\n\n### PyTorch turned the GPU off without saying so\n\nI set up the environment, ran the benchmark, and the first line said `device: cpu`. The card was free and visible. The reason:\n\n`2.13.0+cu130   driver 570.144   torch.cuda.is_available() → False`\nPyTorch had resolved to a build for CUDA 13.0 while the machine’s driver exposes 12.8. Instead of failing, **it silently turns the GPU off** and carries on with the CPU. The warning only shows up if you inspect `torch.cuda.is_available()` by hand.\n\nIt is fixed by pinning the version to the right index:\n\n```\npip install --no-cache-dir \\\n  --index-url https://download.pytorch.org/whl/cu128 torch==2.9.1+cu128\n```\nThis is the second time this year the same trap has cost me a whole run. If a GPU benchmark gives suspiciously slow numbers, that is the first place to look.\n\n### OpenML’s API returned 504 on every endpoint\n\nThe original plan was to take the datasets from OpenML, which is the canonical source for these comparisons. It failed completely during the run:\n\n```\nhttps://api.openml.org/api/v1/json/data/31          → 504 (16.9 s)\nhttps://www.openml.org/api/v1/json/data/31          → 504 (15.7 s)\nhttps://api.openml.org/api/v1/json/data/features/31 → 504 (16.5 s)\n```\nFour retries with growing backoff made no difference: it was not a rate problem, it was the gateway being down. I switched the source to the same benchmark’s CSVs hosted on HuggingFace, which resolved in 250 milliseconds. It is an uncomfortable reminder of how much reproducible research depends on a single service.\n\n### The most-cited model no longer downloads without an account\n\nThis is the finding that interests me most, because it is not technical.\n\nTabPFN is the name that appears in nearly all coverage of this topic. I installed the current version, 8.3.0, and on the first `fit`:\n\n```\nTabPFNLicenseError: TabPFN requires a one-time license acceptance\nto download model weights for local inference, but no interactive\nterminal is available.\n```\nTo fetch the weights you have to open a browser, register, accept the license in a tab on the vendor’s site and export an account token. The best-known tabular foundation model stopped being something you install and run.\n\nThere are two ways out, and I tried both:\n\n- **TabPFN’s 2.x series still downloads the weights without registering.** I installed 2.2.1 in a separate environment and it worked first try. That is the one in the tables below.\n- **TabICL, from Inria’s Soda team, has a three-clause BSD license** and downloads with no barrier at all. It also turned out to be the better of the two.\n\n## The numbers\n\nFourteen datasets, all trimmed to 3,000 rows so the comparison is homogeneous. Accuracy as the median of five seeds. The seconds column sums fitting and prediction, which is each one’s honest cost:\n\n| dataset | rows × cols | TabICL | TabPFN 2.2.1 | XGBoost | Tuned XGB | s TabICL | s tuned XGB | \n|---|---|---|---|---|---|---|---|\n| bank-marketing | 3000 × 7 | 0.7944 | **0.7967** | 0.7711 | 0.7833 | 0.8 | 1.8 | \n| credit | 3000 × 10 | **0.7667** | 0.7578 | 0.7478 | 0.7533 | 0.8 | 2.2 | \n| heloc | 3000 × 22 | 0.7222 | **0.7300** | 0.7022 | 0.7078 | 0.8 | 1.6 | \n| pol | 3000 × 26 | **0.9844** | 0.9833 | 0.9778 | 0.9756 | 0.8 | 1.0 | \n| eye_movements | 3000 × 20 | 0.6100 | **0.6111** | 0.5944 | 0.5833 | 0.8 | 3.0 | \n| california | 3000 × 8 | 0.8967 | **0.8978** | 0.8833 | 0.8767 | 0.6 | 1.6 | \n| house_16H | 3000 × 16 | **0.8811** | 0.8744 | 0.8722 | 0.8711 | 1.1 | 3.0 | \n| MagicTelescope | 3000 × 10 | **0.8778** | 0.8622 | 0.8489 | 0.8456 | 1.3 | 2.4 | \n| electricity | 3000 × 7 | 0.8178 | 0.8111 | 0.8156 | **0.8178** | 0.8 | 2.4 | \n| Diabetes130US | 3000 × 7 | 0.5911 | **0.5933** | 0.5644 | 0.5900 | 0.8 | 1.4 | \n| default-of-credit | 3000 × 20 | 0.6956 | **0.6967** | 0.6878 | 0.6944 | 0.9 | 4.5 | \n| compas-two-years | 3000 × 11 | 0.6733 | **0.6767** | 0.6444 | 0.6700 | 0.6 | 0.7 | \n| Bioresponse | 3000 × 419 | **0.7922** | 0.7567 | 0.7844 | 0.7744 | 6.0 | 27.2 | \n| albert | 3000 × 31 | 0.6500 | **0.6611** | 0.6511 | **0.6611** | 1.8 | 2.6 | \n\nAgainst tuned XGBoost, by accuracy:\n\n- **TabICL: twelve wins, one tie (electricity) and one loss (albert).** Mean difference +0.0106.\n- **TabPFN 2.2.1: eleven wins, one tie (albert) and two losses** (electricity and Bioresponse). Mean difference +0.0075.\n\nAccuracy, however, is a coarse metric: it depends on the threshold and punishes differently depending on how the classes are balanced. Area under the curve is more informative, and there the result gets sharper:\n\n| model | AUC wins | mean difference | \n|---|---|---|\n| TabICL | **14 of 14** | +0.0114 | \n| TabPFN 2.2.1 | 13 of 14 | +0.0089 | \n\nTabICL beats tuned XGBoost **on every dataset without exception**. Even on albert, where it loses on accuracy, it has a better AUC (0.7101 against 0.7094). That detail is exactly why both metrics are worth looking at: accuracy said “loss” where the probability ordering said “win by a hair”.\n\nThat said, the enthusiasm needs calibrating. Winning fourteen of fourteen is a strong signal of **consistency**, not of crushing superiority: the mean difference is one hundredth. Nobody is going to notice that on a dashboard. What does get noticed is the other thing.\n\n## The cost is the reverse of what you expect\n\nLook again at the table’s last two columns. TabICL resolves a dataset in under a second **without tuning anything**. Tuned XGBoost takes between 0.7 and 27.2 seconds searching hyperparameters, and still comes out below.\n\nThat is the real argument, and it is not accuracy. It is that the expensive part of the work, the one that consumes human and machine time, simply disappears.\n\n## The surprise: the advantage does not break with size\n\nThis is where I expected to dismantle the promise. The standard objection to these models is that they only work on toy tables, because the training rows have to fit inside the transformer’s context.\n\nI took jannis, which has 57,580 rows, and trimmed it to growing sizes. Three seeds per size:\n\n| rows | TabICL | XGBoost | Tuned XGB | s TabICL | s tuned XGB | \n|---|---|---|---|---|---|\n| 500 | 0.7533 | **0.7733** | 0.7533 | 0.7 | 3.2 | \n| 1,000 | **0.7733** | 0.7467 | 0.7533 | 0.9 | 4.9 | \n| 2,000 | **0.7800** | 0.7583 | 0.7667 | 1.1 | 9.4 | \n| 4,000 | **0.7817** | 0.7692 | 0.7725 | 1.8 | 21.1 | \n| 8,000 | **0.7958** | 0.7658 | 0.7583 | 2.6 | 19.7 | \n| 16,000 | **0.8090** | 0.7785 | 0.7852 | 5.3 | 34.6 | \n| 32,000 | **0.8230** | 0.7894 | 0.7929 | 11.7 | 50.3 | \n\nThe advantage does not break, and at the large end it consolidates: +0.030 at 32,000 rows. And the split of times opens in the direction opposite to intuition: 11.7 seconds against 50.3.\n\nIt is worth saying precisely what grows and what does not, because the difference matters. What rises steadily is **TabICL’s absolute accuracy**: 0.7533, 0.7733, 0.7800, 0.7817, 0.7958, 0.8090, 0.8230, monotone across all seven measurements. The **advantage over XGBoost**, by contrast, is irregular: 0.000, +0.020, +0.013, +0.009, +0.038, +0.024, +0.030. It rises and falls.\n\nAnd at 4,000 rows that +0.009 advantage is **smaller than the spread across the three seeds** (TabICL ranges from 0.768 to 0.799; tuned XGBoost from 0.758 to 0.790), so at that point it is indistinguishable from noise and I do not count it as a win.\n\nWhere it is solid is at the top: at 32,000 rows TabICL’s **worst** seed (0.8216) sits above XGBoost’s **best** (0.7950). There is no possible overlap there.\n\nThe only place XGBoost wins cleanly is the small end, at 500 rows, where the variability between seeds is so high that I would not bet anything on that difference.\n\nIf you were expecting, as I was, the curve to flip at some point, it does not do so here. You would have to go considerably higher to find it.\n\n## Where they do break\n\nBioresponse is the telltale dataset: 419 columns.\n\nTabPFN 2.2.1 falls to 0.7567, the worst of the four contenders, below even untuned XGBoost. And it costs it 18.1 seconds, twenty-five times more than on a normal dataset. TabICL holds up much better (0.7922, the best of the four) but pays too: 6.0 seconds against the usual 0.8.\n\nThe reading is that the table’s width, not its length, is the axis that genuinely squeezes. That makes sense: the number of columns enters the transformer’s attention cost, and the synthetic pretraining covers tables of tens of columns well, not hundreds.\n\n## What this measurement does not say\n\nI would rather list the limits than pretend they do not exist:\n\n- **Binary classification only.** I did not test regression or multiclass.\n- **A 3,000-row ceiling in the main table.** The scale sweep reaches 32,000, but on a single dataset.\n- **The hyperparameter search was twenty-five combinations.** More aggressive tuning would close part of the gap; how much, I do not know, because I did not measure it.\n- **XGBoost’s search optimized accuracy, and afterwards I also compare by area under the curve.** Which means the “fourteen of fourteen on AUC” is against a boosting model that was not tuned for that metric. Tuning it for AUC would probably improve it there; I did not measure that.\n- **Categoricals were integer-encoded for everyone.** Better treatment would favour XGBoost more than the others.\n- **The encoding and null-filling were computed over the whole table, before splitting.** It is the same leak for all four models, so it does not change who wins, but it inflates everyone slightly and should not be done that way.\n- **Several accuracy wins are smaller than the variation between seeds.**`Diabetes130US` is won by 0.0011 when the per-seed difference ranges from −0.011 to +0.049: there the sign depends on which seed comes up. The area-under-the-curve count does hold seed by seed; the accuracy one, on three or four datasets, does not.\n- **The wide-table finding rests on a single dataset.**`Bioresponse` is the only one with hundreds of columns, so “width is what squeezes” is a hypothesis with one observation, not a rule.\n- **The current version of TabPFN, 8.3.0, went unmeasured** because of the license barrier. TabPFN’s numbers are from 2.2.1, two series behind.\n- **The CPU was shared** with other work on the machine. That affects XGBoost’s times more than the GPU’s, so if anything it plays against the foundation models in the timing comparison.\n\n## When I would use it and when not\n\n**I would use it** on any table between one thousand and thirty thousand rows with fewer than a hundred columns, where a person’s time is worth more than the last hundredth. A second of compute, zero tuning, and a result that in my measurement was consistently better. For an initial exploration it is hard to justify not doing it.\n\n**I would not use it** on very wide tables, where it degrades and gets expensive. Nor in production without a GPU, nor where you need a small artifact that can be inspected and deployed in a lightweight container: a trained tree weighs kilobytes and runs anywhere, whereas here you have to load a transformer.\n\nAnd if the criteria include being able to audit where the weights came from, today the answer is TabICL. Not for performance, though it was also the better of the two, but because it is the one you can still download and run without asking anyone’s permission.\n\n*Measured on 16 August 2026 on a 16 GB RTX 4070 Ti SUPER, with torch 2.9.1+cu128, XGBoost 3.4.1, TabICL 2.1.1 and TabPFN 2.2.1. Data from the `inria-soda/tabular-benchmark` tabular suite. The main table took 585 seconds and the scale sweep 514.*\n\n## Sources\n\n- Accurate predictions on small data with a tabular foundation model, Hollmann et al., *Nature* 637 (2025). The TabPFN v2 paper, with the synthetic-table pretraining method the whole idea comes from.\n- TabPFN: A Transformer That Solves Small Tabular Classification Problems in a Second, arXiv:2207.01848. The original 2022 work, useful for seeing what changed between the first version and the current one.\n- TabICL: A Tabular Foundation Model for In-Context Learning on Large Data, arXiv:2502.05564. The model from Inria’s Soda team that won my measurement, with its proposal for scaling to large tables.\n- soda-inria/tabicl. The implementation, three-clause BSD licensed: it downloads without registering.\n- PriorLabs/TabPFN. TabPFN’s repository, where the license and account-token barrier introduced from series 8 onward is visible.\n\n## Comments\n\nNo comments yet. The first one is yours.","body_html":"<h1 id=\"tabpfn-and-tabicl-against-tuned-xgboost-the-model-that-does-not-\">TabPFN and TabICL against tuned XGBoost: the model that does not train won on fourteen tables out of fourteen</h1>\n<p>The claim behind TabPFN and TabICL is that they predict on a table without ever training on it and still beat tuned boosting. I measured it on fourteen datasets from the Grinsztajn benchmark, with the same split and the same clock for everyone. The one that does not train wins, the advantage holds up to 32,000 rows instead of breaking, and the most-cited model can no longer be downloaded without an account.</p>\n<p>The claim has been going around for months and it is concrete enough to be measurable: a tabular foundation model predicts on a table <strong>without ever having trained on it</strong> and still beats tuned boosting.</p>\n<p>If that is true, half a decade of practice changes shape. Searching hyperparameters stops being a mandatory step and becomes a luxury that sometimes does not pay off.</p>\n<p>So I put it to the test on my own card, with fourteen datasets, four contenders and the same stopwatch for everyone.</p>\n<h2 id=\"what-exactly-these-models-do\">What exactly these models do</h2>\n<p>A tabular foundation model is pretrained on millions of synthetic tables generated on purpose. When a new table arrives, it does not adjust a single weight: it receives the training rows <strong>as context</strong> and produces the predictions in one forward pass.</p>\n<p>It is the same in-context learning idea we already know from language models, moved from words to columns. That is why the verb “train” sits oddly: the code still calls it <code>fit</code>, but inside there is no gradient descent, there is a copy of data to the card.</p>\n<p>That also explains why the cost shows up where you do not expect it. Fitting is nearly free and <strong>prediction</strong> is what pays, exactly the reverse of a tree.</p>\n<p>Put that way it sounds abstract, so here is the full journey of one row: a request arrives with an empty cell, the model weighs it against everything that already happened, resolves it in one pass and returns a future. Pick any of the six cases I measured to see it with their real data:</p>\n<p>The request</p>\n<p>The context</p>\n<p>A single pass</p>\n<p>The prediction</p>\n<p>A credit application arrives</p>\n<p>2 100 applications already settled · 10 columns<code>clf_num/credit.csv</code></p>\n<p>TabICL · nothing tuned · <strong>0.8 s</strong></p>\n<p>Will they repay?</p>\n<p>repaysdefaults</p>\n<p>AUC <strong>0.8528</strong></p>\n<p>credit</p>\n<p>A contact enters the list</p>\n<p>2 100 calls already made · 7 columns<code>clf_num/bank-marketing.csv</code></p>\n<p>TabICL · nothing tuned · <strong>0.8 s</strong></p>\n<p>Is it worth calling them?</p>\n<p>signs updoes not</p>\n<p>AUC <strong>0.8699</strong></p>\n<p>bank-marketing</p>\n<p>A patient is discharged</p>\n<p>2 100 previous discharges · 7 columns<code>clf_num/Diabetes130US.csv</code></p>\n<p>TabICL · nothing tuned · <strong>0.8 s</strong></p>\n<p>Will they be readmitted?</p>\n<p>returnsdoes not</p>\n<p>AUC <strong>0.6484</strong></p>\n<p>Diabetes130US</p>\n<p>A case is assessed</p>\n<p>2 100 cases already closed · 11 columns<code>clf_cat/compas-two-years.csv</code></p>\n<p>TabICL · nothing tuned · <strong>0.6 s</strong></p>\n<p>Will they reoffend?</p>\n<p>reoffendsdoes not</p>\n<p>AUC <strong>0.7329</strong></p>\n<p>compas-two-years</p>\n<p>A market period closes</p>\n<p>2 100 previous periods · 7 columns<code>clf_num/electricity.csv</code></p>\n<p>TabICL · nothing tuned · <strong>0.8 s</strong></p>\n<p>Does the price go up or down?</p>\n<p>updown</p>\n<p>AUC <strong>0.8873</strong></p>\n<p>electricity</p>\n<p>A candidate molecule arrives</p>\n<p>2 100 molecules already assayed · 419 columns<code>clf_num/Bioresponse.csv</code></p>\n<p>TabICL · nothing tuned · <strong>6.0 s</strong></p>\n<p>Does it trigger a biological response?</p>\n<p>activeinert</p>\n<p>AUC <strong>0.8667</strong></p>\n<p>Bioresponse</p>\n<p>The request</p>\n<p>The context</p>\n<p>A single pass</p>\n<p>The prediction</p>\n<p>A credit application arrives</p>\n<p>2 100 applications already settled · 10 columns<code>clf_num/credit.csv</code></p>\n<p>TabICL · nothing tuned · <strong>0.8 s</strong></p>\n<p>Will they repay?</p>\n<p>repaysdefaults</p>\n<p>AUC <strong>0.8528</strong></p>\n<p>credit</p>\n<p>A contact enters the list</p>\n<p>2 100 calls already made · 7 columns<code>clf_num/bank-marketing.csv</code></p>\n<p>TabICL · nothing tuned · <strong>0.8 s</strong></p>\n<p>Is it worth calling them?</p>\n<p>signs updoes not</p>\n<p>AUC <strong>0.8699</strong></p>\n<p>bank-marketing</p>\n<p>A patient is discharged</p>\n<p>2 100 previous discharges · 7 columns<code>clf_num/Diabetes130US.csv</code></p>\n<p>TabICL · nothing tuned · <strong>0.8 s</strong></p>\n<p>Will they be readmitted?</p>\n<p>returnsdoes not</p>\n<p>AUC <strong>0.6484</strong></p>\n<p>Diabetes130US</p>\n<p>A case is assessed</p>\n<p>2 100 cases already closed · 11 columns<code>clf_cat/compas-two-years.csv</code></p>\n<p>TabICL · nothing tuned · <strong>0.6 s</strong></p>\n<p>Will they reoffend?</p>\n<p>reoffendsdoes not</p>\n<p>AUC <strong>0.7329</strong></p>\n<p>compas-two-years</p>\n<p>A market period closes</p>\n<p>2 100 previous periods · 7 columns<code>clf_num/electricity.csv</code></p>\n<p>TabICL · nothing tuned · <strong>0.8 s</strong></p>\n<p>Does the price go up or down?</p>\n<p>updown</p>\n<p>AUC <strong>0.8873</strong></p>\n<p>electricity</p>\n<p>A candidate molecule arrives</p>\n<p>2 100 molecules already assayed · 419 columns<code>clf_num/Bioresponse.csv</code></p>\n<p>TabICL · nothing tuned · <strong>6.0 s</strong></p>\n<p>Does it trigger a biological response?</p>\n<p>activeinert</p>\n<p>AUC <strong>0.8667</strong></p>\n<p>Bioresponse</p>\n<h3 id=\"default-risk\">Default risk</h3>\n<p>the foundation model wins\nBanking’s most repeated case: deciding who gets lent to.</p>\n<div class=\"table-wrap\"><table><thead><tr><th>dataset</th><th>size</th><th>TabICL</th><th>TabPFN</th><th>tuned XGB</th></tr></thead><tbody><tr><td>credit</td><td>3000 × 10</td><td>0.7667</td><td>0.7578</td><td>0.7533</td></tr><tr><td>heloc</td><td>3000 × 22</td><td>0.7222</td><td>0.7300</td><td>0.7078</td></tr><tr><td>default-of-credit</td><td>3000 × 20</td><td>0.6956</td><td>0.6967</td><td>0.6944</td></tr></tbody></table></div>\n<p>All three datasets go to the foundation model. On heloc TabPFN takes 0.022 and TabICL 0.014: in credit risk that is not decoration.</p>\n<p>Context: <code>clf_num/credit.csv</code> from the inria-soda/tabular-benchmark suite, trimmed to 3,000 rows with seed 0 and split 70/30 with stratification.</p>\n<p>sha256 <code>22a759296600c39a56884dafd84eb346be018c537ff096adc8c2ae7e0520f2a9</code></p>\n<h3 id=\"sales-campaign\">Sales campaign</h3>\n<p>the foundation model wins\nWho to call first when there are a thousand contacts and time for a hundred.</p>\n<div class=\"table-wrap\"><table><thead><tr><th>dataset</th><th>size</th><th>TabICL</th><th>TabPFN</th><th>tuned XGB</th></tr></thead><tbody><tr><td>bank-marketing</td><td>3000 × 7</td><td>0.7944</td><td>0.7967</td><td>0.7833</td></tr></tbody></table></div>\n<p>The only case where TabPFN ends up ahead of TabICL. Both beat tuned boosting.</p>\n<p>Context: <code>clf_num/bank-marketing.csv</code> from the inria-soda/tabular-benchmark suite, trimmed to 3,000 rows with seed 0 and split 70/30 with stratification.</p>\n<p>sha256 <code>3433fecd416ce949692442882669881746e8a6b24b904a5f808cdcb196443e7e</code></p>\n<h3 id=\"clinical-readmission\">Clinical readmission</h3>\n<p>the foundation model wins\nA real hospital record: which patient gets readmitted.</p>\n<div class=\"table-wrap\"><table><thead><tr><th>dataset</th><th>size</th><th>TabICL</th><th>TabPFN</th><th>tuned XGB</th></tr></thead><tbody><tr><td>Diabetes130US</td><td>3000 × 7</td><td>0.5911</td><td>0.5933</td><td>0.5900</td></tr></tbody></table></div>\n<p>On accuracy they nearly tie, but the area under the curve opens up sharply: 0.6484 against 0.6264. The risk ordering, which is what a triage uses, improves considerably more than accuracy suggests.</p>\n<p>Context: <code>clf_num/Diabetes130US.csv</code> from the inria-soda/tabular-benchmark suite, trimmed to 3,000 rows with seed 0 and split 70/30 with stratification.</p>\n<p>sha256 <code>9be384ad7edbb9a98b509adb9cc08578e63bd87545b05d97d5147aa15b71385a</code></p>\n<h3 id=\"recidivism\">Recidivism</h3>\n<p>the foundation model wins\nThe dataset that opened the debate on algorithmic bias in the courts.</p>\n<div class=\"table-wrap\"><table><thead><tr><th>dataset</th><th>size</th><th>TabICL</th><th>TabPFN</th><th>tuned XGB</th></tr></thead><tbody><tr><td>compas-two-years</td><td>3000 × 11</td><td>0.6733</td><td>0.6767</td><td>0.6700</td></tr></tbody></table></div>\n<p>The foundation model wins, and it is worth saying that a better model here does not make the use legitimate: this dataset’s argument was never about accuracy.</p>\n<p>Context: <code>clf_cat/compas-two-years.csv</code> from the inria-soda/tabular-benchmark suite, trimmed to 3,000 rows with seed 0 and split 70/30 with stratification.</p>\n<p>sha256 <code>7e0bcef09ae633ad81e26b0a8f8f96dbc4ac3c29b9da0760d38ef2a75adde8d8</code></p>\n<h3 id=\"electricity-demand\">Electricity demand</h3>\n<p>a tie\nConsumption and price series from an electricity market.</p>\n<div class=\"table-wrap\"><table><thead><tr><th>dataset</th><th>size</th><th>TabICL</th><th>TabPFN</th><th>tuned XGB</th></tr></thead><tbody><tr><td>electricity</td><td>3000 × 7</td><td>0.8178</td><td>0.8111</td><td>0.8178</td></tr></tbody></table></div>\n<p>The tie. TabICL matches tuned boosting to the fourth decimal and TabPFN falls below. It is one of the two cases where the advantage does not show up.</p>\n<p>Context: <code>clf_num/electricity.csv</code> from the inria-soda/tabular-benchmark suite, trimmed to 3,000 rows with seed 0 and split 70/30 with stratification.</p>\n<p>sha256 <code>d6005007c4b1f7ba88cf91cb1f231969cce3e49c98d445338c0ab0342ead5d7e</code></p>\n<h3 id=\"molecular-screening\">Molecular screening</h3>\n<p>depends on the model\n419 columns of chemical descriptors: the wide table.</p>\n<div class=\"table-wrap\"><table><thead><tr><th>dataset</th><th>size</th><th>TabICL</th><th>TabPFN</th><th>tuned XGB</th></tr></thead><tbody><tr><td>Bioresponse</td><td>3000 × 419</td><td>0.7922</td><td>0.7567</td><td>0.7744</td></tr></tbody></table></div>\n<p>This is where TabPFN breaks: 0.7567 is worse than even untuned XGBoost, and it costs 18.1 seconds. TabICL holds up and comes first. The table’s width, not its length, is what squeezes.</p>\n<p>Context: <code>clf_num/Bioresponse.csv</code> from the inria-soda/tabular-benchmark suite, trimmed to 3,000 rows with seed 0 and split 70/30 with stratification.</p>\n<p>sha256 <code>bd3d277821eb949df41219363f0386549019249776fccf7a9f08d3bec10b8727</code></p>\n<h2 id=\"what-this-is-for-concretely\">What this is for, concretely</h2>\n<p>The six cases in the diagram are not brochure examples: they are the datasets I measured with. But it is worth widening the map, because “tabular foundation model” sounds like a laboratory and the problem it solves is one of the most common there is.</p>\n<p>A table is a spreadsheet: rows that are cases and columns that are attributes. And the task is always the same, <strong>predict one column from the others</strong>:</p>\n<ul><li>A customer table with tenure, plan, usage and complaints, to estimate <strong>which ones are going to leave next month</strong> .</li><li>A transaction log with amount, merchant, hour and country, to flag <strong>which ones are fraud</strong> .</li><li>A history of credit applications with income, debt and payment behaviour, to estimate <strong>who is not going to repay</strong> . Three of the datasets I measured are exactly that:<code>credit</code> ,<code>heloc</code> and<code>default-of-credit</code> .</li><li>Sensor readings from a machine, to anticipate <strong>when it is going to break</strong> .</li><li>Patient records with symptoms and lab results, to prioritize <strong>who gets seen first</strong> .<code>Diabetes130US</code> , another of the datasets, is a real hospital record.</li></ul>\n<p>That is half the data work done in any company. It is also the ground where deep learning had been losing for years: for tables, a good gradient-boosted tree ensemble was still the right answer, and the research that assembled these datasets was titled exactly that way, asking why trees still won.</p>\n<h3 id=\"what-changes-if-the-promise-holds\">What changes if the promise holds</h3>\n<p>Today, solving any of those cases has a ritual: prepare the data, choose a model, <strong>search hyperparameters</strong>, validate, repeat. The search is the boring part and the one that eats machine and human hours. In this benchmark, XGBoost’s search took up to 50 seconds per dataset; on a real problem with more rows and more combinations, it is minutes or hours.</p>\n<p>A tabular foundation model proposes skipping that ritual entirely. You hand it the table, it answers in a second, and that is that. No tree depth to choose, no learning rate, no cross-validation to decide among twenty-five candidates.</p>\n<p>Put in concrete terms: instead of spending the afternoon tuning a churn model, you have an answer in the time it takes to make coffee, and only then do you decide whether anything is worth refining. For exploring a table that just arrived, or for having an honest baseline before investing time, it is hard to beat.</p>\n<p>The question, then, is not whether the idea is attractive. It is whether the result holds up when measured.</p>\n<h2 id=\"the-benchmark-s-rules\">The benchmark’s rules</h2>\n<p>A badly built benchmark says whatever you want. These are the rules I imposed on myself before seeing a single number:</p>\n<p><strong>The data is third-party and from the field where this argument is fought.</strong> I used Grinsztajn’s tabular benchmark, the suite gathered for the work asking why trees still beat deep learning on tables. Choosing the datasets yourself is the easiest way to manufacture the result you like.</p>\n<p><strong>XGBoost goes in tuned, not for decoration.</strong> The claim says “tuned boosting”, so comparing against factory parameters would be a straw man. I ran two versions: one with fixed, reasonable parameters and another with a random search over twenty-five combinations and three-fold cross-validation.</p>\n<p><strong>Same split, same seed, same columns for everyone.</strong> Categoricals are integer-encoded. It is not the best possible treatment, but it is identical for all four, which is what makes the number comparable.</p>\n<p><strong>Five seeds per dataset, and the median is reported.</strong> A single split cannot tell signal from luck.</p>\n<p><strong>Prediction time includes two inference passes</strong>, the class one and the probability one, because the benchmark needs both to compute accuracy and area under the curve. That makes things more expensive precisely for the foundation models, which is where their cost lives, so the seconds I publish for them are inflated in nobody’s favour.</p>\n<p><strong>Two cores are left free.</strong> The machine has other work on it, and a benchmark that eats the whole CPU measures the fight with the scheduler, not the model.</p>\n<p>Everything ran on a 16 GB RTX 4070 Ti SUPER, with fourteen cores for XGBoost.</p>\n<h2 id=\"three-stumbles-before-the-first-number\">Three stumbles before the first number</h2>\n<p>Publishing only the final table would be lying by omission. This is what it cost to get there.</p>\n<h3 id=\"pytorch-turned-the-gpu-off-without-saying-so\">PyTorch turned the GPU off without saying so</h3>\n<p>I set up the environment, ran the benchmark, and the first line said <code>device: cpu</code>. The card was free and visible. The reason:</p>\n<p><code>2.13.0+cu130   driver 570.144   torch.cuda.is_available() → False</code>\nPyTorch had resolved to a build for CUDA 13.0 while the machine’s driver exposes 12.8. Instead of failing, <strong>it silently turns the GPU off</strong> and carries on with the CPU. The warning only shows up if you inspect <code>torch.cuda.is_available()</code> by hand.</p>\n<p>It is fixed by pinning the version to the right index:</p>\n<pre><code>pip install --no-cache-dir \\\n  --index-url https://download.pytorch.org/whl/cu128 torch==2.9.1+cu128</code></pre>\n<p>This is the second time this year the same trap has cost me a whole run. If a GPU benchmark gives suspiciously slow numbers, that is the first place to look.</p>\n<h3 id=\"openml-s-api-returned-504-on-every-endpoint\">OpenML’s API returned 504 on every endpoint</h3>\n<p>The original plan was to take the datasets from OpenML, which is the canonical source for these comparisons. It failed completely during the run:</p>\n<pre><code>https://api.openml.org/api/v1/json/data/31          → 504 (16.9 s)\nhttps://www.openml.org/api/v1/json/data/31          → 504 (15.7 s)\nhttps://api.openml.org/api/v1/json/data/features/31 → 504 (16.5 s)</code></pre>\n<p>Four retries with growing backoff made no difference: it was not a rate problem, it was the gateway being down. I switched the source to the same benchmark’s CSVs hosted on HuggingFace, which resolved in 250 milliseconds. It is an uncomfortable reminder of how much reproducible research depends on a single service.</p>\n<h3 id=\"the-most-cited-model-no-longer-downloads-without-an-account\">The most-cited model no longer downloads without an account</h3>\n<p>This is the finding that interests me most, because it is not technical.</p>\n<p>TabPFN is the name that appears in nearly all coverage of this topic. I installed the current version, 8.3.0, and on the first <code>fit</code>:</p>\n<pre><code>TabPFNLicenseError: TabPFN requires a one-time license acceptance\nto download model weights for local inference, but no interactive\nterminal is available.</code></pre>\n<p>To fetch the weights you have to open a browser, register, accept the license in a tab on the vendor’s site and export an account token. The best-known tabular foundation model stopped being something you install and run.</p>\n<p>There are two ways out, and I tried both:</p>\n<ul><li><strong>TabPFN’s 2.x series still downloads the weights without registering.</strong> I installed 2.2.1 in a separate environment and it worked first try. That is the one in the tables below.</li><li><strong>TabICL, from Inria’s Soda team, has a three-clause BSD license</strong> and downloads with no barrier at all. It also turned out to be the better of the two.</li></ul>\n<h2 id=\"the-numbers\">The numbers</h2>\n<p>Fourteen datasets, all trimmed to 3,000 rows so the comparison is homogeneous. Accuracy as the median of five seeds. The seconds column sums fitting and prediction, which is each one’s honest cost:</p>\n<div class=\"table-wrap\"><table><thead><tr><th>dataset</th><th>rows × cols</th><th>TabICL</th><th>TabPFN 2.2.1</th><th>XGBoost</th><th>Tuned XGB</th><th>s TabICL</th><th>s tuned XGB</th></tr></thead><tbody><tr><td>bank-marketing</td><td>3000 × 7</td><td>0.7944</td><td><strong>0.7967</strong></td><td>0.7711</td><td>0.7833</td><td>0.8</td><td>1.8</td></tr><tr><td>credit</td><td>3000 × 10</td><td><strong>0.7667</strong></td><td>0.7578</td><td>0.7478</td><td>0.7533</td><td>0.8</td><td>2.2</td></tr><tr><td>heloc</td><td>3000 × 22</td><td>0.7222</td><td><strong>0.7300</strong></td><td>0.7022</td><td>0.7078</td><td>0.8</td><td>1.6</td></tr><tr><td>pol</td><td>3000 × 26</td><td><strong>0.9844</strong></td><td>0.9833</td><td>0.9778</td><td>0.9756</td><td>0.8</td><td>1.0</td></tr><tr><td>eye_movements</td><td>3000 × 20</td><td>0.6100</td><td><strong>0.6111</strong></td><td>0.5944</td><td>0.5833</td><td>0.8</td><td>3.0</td></tr><tr><td>california</td><td>3000 × 8</td><td>0.8967</td><td><strong>0.8978</strong></td><td>0.8833</td><td>0.8767</td><td>0.6</td><td>1.6</td></tr><tr><td>house_16H</td><td>3000 × 16</td><td><strong>0.8811</strong></td><td>0.8744</td><td>0.8722</td><td>0.8711</td><td>1.1</td><td>3.0</td></tr><tr><td>MagicTelescope</td><td>3000 × 10</td><td><strong>0.8778</strong></td><td>0.8622</td><td>0.8489</td><td>0.8456</td><td>1.3</td><td>2.4</td></tr><tr><td>electricity</td><td>3000 × 7</td><td>0.8178</td><td>0.8111</td><td>0.8156</td><td><strong>0.8178</strong></td><td>0.8</td><td>2.4</td></tr><tr><td>Diabetes130US</td><td>3000 × 7</td><td>0.5911</td><td><strong>0.5933</strong></td><td>0.5644</td><td>0.5900</td><td>0.8</td><td>1.4</td></tr><tr><td>default-of-credit</td><td>3000 × 20</td><td>0.6956</td><td><strong>0.6967</strong></td><td>0.6878</td><td>0.6944</td><td>0.9</td><td>4.5</td></tr><tr><td>compas-two-years</td><td>3000 × 11</td><td>0.6733</td><td><strong>0.6767</strong></td><td>0.6444</td><td>0.6700</td><td>0.6</td><td>0.7</td></tr><tr><td>Bioresponse</td><td>3000 × 419</td><td><strong>0.7922</strong></td><td>0.7567</td><td>0.7844</td><td>0.7744</td><td>6.0</td><td>27.2</td></tr><tr><td>albert</td><td>3000 × 31</td><td>0.6500</td><td><strong>0.6611</strong></td><td>0.6511</td><td><strong>0.6611</strong></td><td>1.8</td><td>2.6</td></tr></tbody></table></div>\n<p>Against tuned XGBoost, by accuracy:</p>\n<ul><li><strong>TabICL: twelve wins, one tie (electricity) and one loss (albert).</strong> Mean difference +0.0106.</li><li><strong>TabPFN 2.2.1: eleven wins, one tie (albert) and two losses</strong> (electricity and Bioresponse). Mean difference +0.0075.</li></ul>\n<p>Accuracy, however, is a coarse metric: it depends on the threshold and punishes differently depending on how the classes are balanced. Area under the curve is more informative, and there the result gets sharper:</p>\n<div class=\"table-wrap\"><table><thead><tr><th>model</th><th>AUC wins</th><th>mean difference</th></tr></thead><tbody><tr><td>TabICL</td><td><strong>14 of 14</strong></td><td>+0.0114</td></tr><tr><td>TabPFN 2.2.1</td><td>13 of 14</td><td>+0.0089</td></tr></tbody></table></div>\n<p>TabICL beats tuned XGBoost <strong>on every dataset without exception</strong>. Even on albert, where it loses on accuracy, it has a better AUC (0.7101 against 0.7094). That detail is exactly why both metrics are worth looking at: accuracy said “loss” where the probability ordering said “win by a hair”.</p>\n<p>That said, the enthusiasm needs calibrating. Winning fourteen of fourteen is a strong signal of <strong>consistency</strong>, not of crushing superiority: the mean difference is one hundredth. Nobody is going to notice that on a dashboard. What does get noticed is the other thing.</p>\n<h2 id=\"the-cost-is-the-reverse-of-what-you-expect\">The cost is the reverse of what you expect</h2>\n<p>Look again at the table’s last two columns. TabICL resolves a dataset in under a second <strong>without tuning anything</strong>. Tuned XGBoost takes between 0.7 and 27.2 seconds searching hyperparameters, and still comes out below.</p>\n<p>That is the real argument, and it is not accuracy. It is that the expensive part of the work, the one that consumes human and machine time, simply disappears.</p>\n<h2 id=\"the-surprise-the-advantage-does-not-break-with-size\">The surprise: the advantage does not break with size</h2>\n<p>This is where I expected to dismantle the promise. The standard objection to these models is that they only work on toy tables, because the training rows have to fit inside the transformer’s context.</p>\n<p>I took jannis, which has 57,580 rows, and trimmed it to growing sizes. Three seeds per size:</p>\n<div class=\"table-wrap\"><table><thead><tr><th>rows</th><th>TabICL</th><th>XGBoost</th><th>Tuned XGB</th><th>s TabICL</th><th>s tuned XGB</th></tr></thead><tbody><tr><td>500</td><td>0.7533</td><td><strong>0.7733</strong></td><td>0.7533</td><td>0.7</td><td>3.2</td></tr><tr><td>1,000</td><td><strong>0.7733</strong></td><td>0.7467</td><td>0.7533</td><td>0.9</td><td>4.9</td></tr><tr><td>2,000</td><td><strong>0.7800</strong></td><td>0.7583</td><td>0.7667</td><td>1.1</td><td>9.4</td></tr><tr><td>4,000</td><td><strong>0.7817</strong></td><td>0.7692</td><td>0.7725</td><td>1.8</td><td>21.1</td></tr><tr><td>8,000</td><td><strong>0.7958</strong></td><td>0.7658</td><td>0.7583</td><td>2.6</td><td>19.7</td></tr><tr><td>16,000</td><td><strong>0.8090</strong></td><td>0.7785</td><td>0.7852</td><td>5.3</td><td>34.6</td></tr><tr><td>32,000</td><td><strong>0.8230</strong></td><td>0.7894</td><td>0.7929</td><td>11.7</td><td>50.3</td></tr></tbody></table></div>\n<p>The advantage does not break, and at the large end it consolidates: +0.030 at 32,000 rows. And the split of times opens in the direction opposite to intuition: 11.7 seconds against 50.3.</p>\n<p>It is worth saying precisely what grows and what does not, because the difference matters. What rises steadily is <strong>TabICL’s absolute accuracy</strong>: 0.7533, 0.7733, 0.7800, 0.7817, 0.7958, 0.8090, 0.8230, monotone across all seven measurements. The <strong>advantage over XGBoost</strong>, by contrast, is irregular: 0.000, +0.020, +0.013, +0.009, +0.038, +0.024, +0.030. It rises and falls.</p>\n<p>And at 4,000 rows that +0.009 advantage is <strong>smaller than the spread across the three seeds</strong> (TabICL ranges from 0.768 to 0.799; tuned XGBoost from 0.758 to 0.790), so at that point it is indistinguishable from noise and I do not count it as a win.</p>\n<p>Where it is solid is at the top: at 32,000 rows TabICL’s <strong>worst</strong> seed (0.8216) sits above XGBoost’s <strong>best</strong> (0.7950). There is no possible overlap there.</p>\n<p>The only place XGBoost wins cleanly is the small end, at 500 rows, where the variability between seeds is so high that I would not bet anything on that difference.</p>\n<p>If you were expecting, as I was, the curve to flip at some point, it does not do so here. You would have to go considerably higher to find it.</p>\n<h2 id=\"where-they-do-break\">Where they do break</h2>\n<p>Bioresponse is the telltale dataset: 419 columns.</p>\n<p>TabPFN 2.2.1 falls to 0.7567, the worst of the four contenders, below even untuned XGBoost. And it costs it 18.1 seconds, twenty-five times more than on a normal dataset. TabICL holds up much better (0.7922, the best of the four) but pays too: 6.0 seconds against the usual 0.8.</p>\n<p>The reading is that the table’s width, not its length, is the axis that genuinely squeezes. That makes sense: the number of columns enters the transformer’s attention cost, and the synthetic pretraining covers tables of tens of columns well, not hundreds.</p>\n<h2 id=\"what-this-measurement-does-not-say\">What this measurement does not say</h2>\n<p>I would rather list the limits than pretend they do not exist:</p>\n<ul><li><strong>Binary classification only.</strong> I did not test regression or multiclass.</li><li><strong>A 3,000-row ceiling in the main table.</strong> The scale sweep reaches 32,000, but on a single dataset.</li><li><strong>The hyperparameter search was twenty-five combinations.</strong> More aggressive tuning would close part of the gap; how much, I do not know, because I did not measure it.</li><li><strong>XGBoost’s search optimized accuracy, and afterwards I also compare by area under the curve.</strong> Which means the “fourteen of fourteen on AUC” is against a boosting model that was not tuned for that metric. Tuning it for AUC would probably improve it there; I did not measure that.</li><li><strong>Categoricals were integer-encoded for everyone.</strong> Better treatment would favour XGBoost more than the others.</li><li><strong>The encoding and null-filling were computed over the whole table, before splitting.</strong> It is the same leak for all four models, so it does not change who wins, but it inflates everyone slightly and should not be done that way.</li><li><strong>Several accuracy wins are smaller than the variation between seeds.</strong><code>Diabetes130US</code> is won by 0.0011 when the per-seed difference ranges from −0.011 to +0.049: there the sign depends on which seed comes up. The area-under-the-curve count does hold seed by seed; the accuracy one, on three or four datasets, does not.</li><li><strong>The wide-table finding rests on a single dataset.</strong><code>Bioresponse</code> is the only one with hundreds of columns, so “width is what squeezes” is a hypothesis with one observation, not a rule.</li><li><strong>The current version of TabPFN, 8.3.0, went unmeasured</strong> because of the license barrier. TabPFN’s numbers are from 2.2.1, two series behind.</li><li><strong>The CPU was shared</strong> with other work on the machine. That affects XGBoost’s times more than the GPU’s, so if anything it plays against the foundation models in the timing comparison.</li></ul>\n<h2 id=\"when-i-would-use-it-and-when-not\">When I would use it and when not</h2>\n<p><strong>I would use it</strong> on any table between one thousand and thirty thousand rows with fewer than a hundred columns, where a person’s time is worth more than the last hundredth. A second of compute, zero tuning, and a result that in my measurement was consistently better. For an initial exploration it is hard to justify not doing it.</p>\n<p><strong>I would not use it</strong> on very wide tables, where it degrades and gets expensive. Nor in production without a GPU, nor where you need a small artifact that can be inspected and deployed in a lightweight container: a trained tree weighs kilobytes and runs anywhere, whereas here you have to load a transformer.</p>\n<p>And if the criteria include being able to audit where the weights came from, today the answer is TabICL. Not for performance, though it was also the better of the two, but because it is the one you can still download and run without asking anyone’s permission.</p>\n<p><em>Measured on 16 August 2026 on a 16 GB RTX 4070 Ti SUPER, with torch 2.9.1+cu128, XGBoost 3.4.1, TabICL 2.1.1 and TabPFN 2.2.1. Data from the <code>inria-soda/tabular-benchmark</code> tabular suite. The main table took 585 seconds and the scale sweep 514.</em></p>\n<h2 id=\"sources\">Sources</h2>\n<ul><li>Accurate predictions on small data with a tabular foundation model, Hollmann et al., <em>Nature</em> 637 (2025). The TabPFN v2 paper, with the synthetic-table pretraining method the whole idea comes from.</li><li>TabPFN: A Transformer That Solves Small Tabular Classification Problems in a Second, arXiv:2207.01848. The original 2022 work, useful for seeing what changed between the first version and the current one.</li><li>TabICL: A Tabular Foundation Model for In-Context Learning on Large Data, arXiv:2502.05564. The model from Inria’s Soda team that won my measurement, with its proposal for scaling to large tables.</li><li>soda-inria/tabicl. The implementation, three-clause BSD licensed: it downloads without registering.</li><li>PriorLabs/TabPFN. TabPFN’s repository, where the license and account-token barrier introduced from series 8 onward is visible.</li></ul>\n<h2 id=\"comments\">Comments</h2>\n<p>No comments yet. The first one is yours.</p>","headings":[{"level":1,"text":"TabPFN and TabICL against tuned XGBoost: the model that does not train won on fourteen tables out of fourteen","id":"tabpfn-and-tabicl-against-tuned-xgboost-the-model-that-does-not-"},{"level":2,"text":"What exactly these models do","id":"what-exactly-these-models-do"},{"level":3,"text":"Default risk","id":"default-risk"},{"level":3,"text":"Sales campaign","id":"sales-campaign"},{"level":3,"text":"Clinical readmission","id":"clinical-readmission"},{"level":3,"text":"Recidivism","id":"recidivism"},{"level":3,"text":"Electricity demand","id":"electricity-demand"},{"level":3,"text":"Molecular screening","id":"molecular-screening"},{"level":2,"text":"What this is for, concretely","id":"what-this-is-for-concretely"},{"level":3,"text":"What changes if the promise holds","id":"what-changes-if-the-promise-holds"},{"level":2,"text":"The benchmark’s rules","id":"the-benchmark-s-rules"},{"level":2,"text":"Three stumbles before the first number","id":"three-stumbles-before-the-first-number"},{"level":3,"text":"PyTorch turned the GPU off without saying so","id":"pytorch-turned-the-gpu-off-without-saying-so"},{"level":3,"text":"OpenML’s API returned 504 on every endpoint","id":"openml-s-api-returned-504-on-every-endpoint"},{"level":3,"text":"The most-cited model no longer downloads without an account","id":"the-most-cited-model-no-longer-downloads-without-an-account"},{"level":2,"text":"The numbers","id":"the-numbers"},{"level":2,"text":"The cost is the reverse of what you expect","id":"the-cost-is-the-reverse-of-what-you-expect"},{"level":2,"text":"The surprise: the advantage does not break with size","id":"the-surprise-the-advantage-does-not-break-with-size"},{"level":2,"text":"Where they do break","id":"where-they-do-break"},{"level":2,"text":"What this measurement does not say","id":"what-this-measurement-does-not-say"},{"level":2,"text":"When I would use it and when not","id":"when-i-would-use-it-and-when-not"},{"level":2,"text":"Sources","id":"sources"},{"level":2,"text":"Comments","id":"comments"}]}}