Robustness is important: Limitations of LLMs for predictions on tabular data
Clicks: 1
ID: 315432
2026
Article Quality & Performance Metrics
Overall Quality
Not rated
Combines reader engagement with the AI quality analysis. This
article has not been analysed, so there is no overall score —
reader engagement is measured and shown alongside.
Reader Engagement
0.0
/100
1 views
0 readers
AI Quality Assessment
Not analyzed
Readership in this journal
Ranked #67 of 88 articles by views in PNAS nexus
Most read
Least read
Bar heights use a square-root scale.
Mint this article as an NFT
Not yet mintedCreate a permanent, verifiable on-chain record of this article on the Scimatic Network. The NFT is held in your Journament account, and you can withdraw it to your own wallet at any time.
5
SUSD
one-off · no wallet required
Abstract
Abstract Large Language Models (LLMs) are being applied in a wide array of settings, well beyond typical language-oriented use cases. In particular, LLMs are increasingly used as a plug-and-play method for generating predictions on tabular data. Prior work has shown that LLMs, via in-context learning or supervised fine-tuning, perform comparably with many tabular supervised learning techniques. However, we identify a critical vulnerability of using LLMs for tabular prediction -- making changes to data representation that are completely irrelevant to the underlying learning task can drastically alter LLMs' predictions on the same data. For example, simply changing variable names can sway the size of prediction error by as much as 82% in certain settings. Such prediction sensitivity with respect to task-irrelevant variations manifests under both in-context learning and supervised fine-tuning, for both close-weight and open-weight general-purpose LLMs. Moreover, by examining the attention scores of two open-weight LLMs, we discover a non-uniform attention pattern: training examples and variable names/values occupying certain positions in the prompt receive more attention when generating output tokens, even though fundamentally there should not be different emphasis a priori on data rows / columns in specific positions. This partially explains the sensitivity due to task-irrelevant variations. We also consider several state-of-the-art tabular foundation models trained specifically for tabular prediction. They achieve better prediction performance than general-purpose LLMs but are still not immune to task-irrelevant variations. Overall, LLMs (especially general-purpose models) currently lack a basic level of robustness to be used as a principled prediction tool.
| Reference Key |
openalex_W7163042746
Use this key to autocite in the manuscript while using
SciMatic Manuscript Manager or Thesis Manager
|
|---|---|
| Authors | Hejia Liu, Mochen Yang, Gediminas Adomavičius |
| Journal | PNAS nexus |
| Year | 2026 |
| DOI |
10.1093/pnasnexus/pgag197
|
| URL | |
| Keywords | Keywords not found |
Citations
No citations found. To add a citation, contact the admin at info@scimatic.org
Comments
No comments yet. Be the first to comment on this article.