Short answer: classical models remain preferable when the input is a table with a stable schema, the target variable is numerical or categorical, and quality must be confirmed on future periods. Typical examples include estimating the probability of default, forecasting cash flow, ranking counterparties, detecting anomalous transactions, and forecasting a time series from lagged features.
In such tasks, linear and statistical models, Random Forest, and especially gradient boosting over decision trees, or GBDT, provide a strong starting point and often become the final choice. LLMs are more sensible where a substantial part of the task is expressed in natural language: extracting facts from notes to financial statements, converting a question into SQL, combining a table with documents, or explaining an already calculated result.
This does not mean that trees universally outperform LLMs. There are still insufficient direct, independent comparisons of modern LLMs and well-tuned GBDT on a representative set of real financial tables. The practical conclusion is already clear: LLMs should not be treated as an automatic replacement for a specialized tabular pipeline.
What exactly is being compared
Structured financial data here means records with a defined schema: rows of observations and columns of numerical, categorical, or temporal features. These may include financial-statement ratios, payment history, instrument type, industry, currency, publication date, and a target variable such as a loss or an instance of delinquency.
The term “classical models” brings together different tools that should not be considered interchangeable:
- linear and logistic regression are useful as simple, regularized, and comparatively transparent baselines;
- ARIMA, ETS, and related statistical methods are intended primarily for time series;
- Random Forest and GBDT work with tabular features and model nonlinearities and interactions;
- specialized neural networks for tables or sequences belong to deep learning, but are not necessarily LLMs;
- an LLM processes a table as a tokenized representation – for example, text, JSON, Markdown, or a sequence of “column: value” pairs.
The last distinction is fundamental. For GBDT, a number in a cell is a numerical feature. For an LLM, it is initially a sequence of tokens if the model does not call a separate calculator or executable code.
Tabular classification and regression with limited samples
If the task consists of mapping a fixed set of features to a probability, class, or number, GBDT should usually be tested before an LLM. Trees naturally model thresholds, missing values, and nonlinear interactions between financial indicators. They do not need to turn every row into a text context to do so.
On medium-sized tabular datasets, tree-based methods often outperform neural networks even without considering speed. One well-known benchmark considered datasets with approximately 10 000 observations and found a persistent advantage for trees in this setting, while also showing that the result depends on the characteristic properties of tabular data Why do tree-based models still outperform deep learning on tabular data?. This is not evidence for every financial dataset: the conclusion cannot be transferred without verification to millions of observations, multimodal data, or tasks with a large amount of text.
A review of LLM applications to tables also shows a heterogeneous picture: the result depends on the model, table serialization, prompt, fine-tuning, retrieval, and the specific dataset. In some experiments, a simpler KNN with feature weights from XGBoost outperformed an LLM-based approach Large Language Models on Tabular Data – A Survey. Therefore, for ordinary feature-based classification or regression, any advantage for LLMs must be demonstrated rather than assumed.
A classical model is especially attractive when several conditions hold simultaneously:
- there are few or a moderate number of labeled observations;
- the feature schema is stable and known in advance;
- batch processing and low latency are important;
- probabilities need to be estimated rather than text generated;
- a repeatable pipeline with controlled versions of data and features is required;
- the cost of errors differs substantially between false-positive and false-negative decisions.
However, “classical” does not mean “linear only.” In an empirical study of return forecasting, trees and neural networks accounted for nonlinear feature interactions and improved forecasts compared with simpler regression specifications Empirical Asset Pricing via Machine Learning. This study did not compare models with LLMs, but it clearly demonstrates the breadth of specialized tools.
Precise arithmetic should not depend on generation
Calculating a liquidity ratio, summing payments, sorting by return, applying a threshold rule, and constructing an accounting reconciliation are not tasks for probabilistic generation. If the formula is known, the appropriate tool is usually SQL, Python, a spreadsheet, or a verifiable computational graph.
For example, with current assets of 120 million and current liabilities of 80 million, the current ratio is 1.5. No trained model is needed to obtain this result. An LLM may recognize the request and generate the expression 120 / 80, but the division itself is better performed by a deterministic tool.
Tokenization of numbers does not provide a language model with reliable algorithmic understanding of arithmetic. A review of tabular LLMs identifies numerical reasoning, table scale, and dependence on the representation method as substantial limitations of such systems Large Language Models on Tabular Data – A Survey. Calling a calculator or Python materially changes the solution architecture: the comparison should then be between two complete systems, including their tools, checks, and error handling, rather than between “LLM and GBDT.”
Financial time series require a temporal evaluation model
For forecasting price, return, volume, volatility, or cash flow, the order of observations is part of the task. A reasonable set of baselines may include a naive forecast, ARIMA or ETS, regularized regression, and GBDT with lags and rolling statistics. A Transformer or LLM makes sense as an addition after them, not as a replacement for them.
In one published comparison using daily data for six US stocks over 2014–2024, ARIMA and Random Forest remained competitive, while LSTM and Transformer did not show a stable advantage for all tickers A Comparative Study of Transformer-Based and Classical Models for Financial Time-Series Forecasting. The scope of this conclusion is narrow: the horizon was one trading day, and the study does not establish the universal superiority of ARIMA or Random Forest.
The evaluation method may matter much more than the name of the architecture. Data leakage, or leakage, occurs when information that was unavailable at the time of the forecasted decision is used during training or feature construction. In financial data, this may be a financial statement tied to the end of a quarter but published later, a revised macroeconomic indicator, or normalization calculated over the entire history, including the test period.
Randomly shuffling rows usually violates the temporal boundary and can inflate a backtest estimate. Studies of leakage in financial testing show that even common data-preparation procedures can pass information from the future to the model Information Leakage in Backtesting.
Instead of a random split, use a temporal split or walk-forward validation: the model is trained on the available past, tested on the next time interval, and then the boundary is moved forward. This protocol does not eliminate survivorship bias, publication delays, data revisions, multiple testing, or operating costs, but it better reproduces the real decision-making point.
Calibration and robustness matter more than a single metric
A high AUC or low RMSE does not by itself mean that a model is suitable for a financial decision.
Calibration shows whether the predicted probability corresponds to the observed event frequency. If default occurs in approximately one out of ten applications with a predicted default probability of 10%, the model is calibrated in that range. For limits, reserves, and decisions with asymmetric error costs, this may matter more than a small increase in a ranking metric.
Consider a conditional, non-empirical example. On a walk-forward test, logistic regression received an AUROC of 0.842, GBDT – 0.867, and an LLM predictor – 0.869. However, after calibration, the Brier score for GBDT was 0.076, while that for the LLM was 0.091; in addition, the LLM did not fit within the established latency budget. Choosing the LLM solely because of the 0.002 increase in AUROC would be unjustified: the system estimates absolute risk less accurately and does not meet the operational constraint. The numbers here merely illustrate the selection logic.
Concept drift – a change over time in the relationship between features and the target variable – must also be checked. For example, the same debt ratio may have a different relationship with risk after a change in interest rates or lending standards. Drift affects both classical models and LLMs; no architecture eliminates the need for monitoring across time periods.
Auditability, reproducibility, and schema stability
Interpretability should mean the ability to trace why a model produced a result and which inputs influenced it. Linear-model coefficients, feature-transformation rules, and feature attribution for GBDT are usually easier to include in a controlled report than freely generated LLM text. But attribution shows model behavior, not a causal effect of a feature.
An LLM adds further sources of variability: prompt template, column order, how missing values are marked, number format, model version, and generation parameters. In experiments on table understanding, quality changed when rows and columns were permuted, when tables were transposed, and when the format was changed while the same information was retained Rethinking Tabular Data Understanding with Large Language Models. These results were obtained mainly on table-understanding tasks and do not prove the same degradation for every financial dataset. They nevertheless show that serialization is part of the model and must be tested as code.
Classical models are not automatically deterministic either: the result may change with the seed, parallel training, library version, or data order. The difference is that such a pipeline is usually easier to fix, reproduce, and cover with tests.
| Requirement | When the classical pipeline more often has the advantage | What is required from an LLM system |
|---|---|---|
| Precise arithmetic | Formulas are executed directly in SQL or code | A call to a verifiable tool rather than number generation |
| Fixed tabular schema | GBDT or regression accepts typed features | Stable serialization and tests for order, format, and missing values |
| Calibrated risk | Standard calibration and diagnostic procedures are available | A separate probabilistic output and its temporal validation |
| Low latency and high-volume inference | A compact model processes rows in batches | Measurements of tokens, caching, and queues are needed |
| Audit | Features, parameters, and calculations are easier to version | The prompt, context, model version, and tool calls must be retained |
| Working with financial-reporting text | A separate NLP layer is required | The LLM can be the main part of the interface or extraction |
Where LLMs genuinely add value
The strong area for LLMs begins at the boundary between tables and language. A model can:
- convert an analyst’s question into SQL;
- match wording from a note to financial statements with a column in an analytical data mart;
- extract candidate facts from a filing for subsequent verification;
- explain a GBDT result in terms of the original indicators;
- orchestrate queries to a database, Python code, and document search;
- combine a numerical table with descriptions of risks, contracts, or management commentary.
A practical architecture separates responsibilities. The LLM interprets intent and forms a plan; SQL or Python performs the calculation; a specialized model produces the forecast; and a verifiable layer assembles data provenance and the result.
This separation is especially important for regulatory reporting. For example, the SEC publishes Financial Statement Data Sets with numerical data extracted from XBRL and converted into a flat format for analysis across companies and periods. At the same time, the SEC warns that the datasets do not replace the original filings, may contain presentation errors, and do not include all available metadata Financial Statement Data Sets – SEC.gov. An LLM can help find and explain a discrepancy, but it should not turn an unverified extraction into an authoritative number.
A practical selection process
The comparison should begin not with the architecture, but with the decision point and the information available at that time.
- Fix the target variable and temporal boundary. For each feature, the date when it was actually available must be known, not just the reporting period.
- Separate calculation from prediction. Known formulas are executed deterministically; a model is used only where an unknown relationship must be estimated.
- Build simple baselines. A naive forecast, regularized regression, specialized time-series model, and GBDT establish the minimum level that a more complex system must exceed.
- Conduct walk-forward evaluation. Compare quality across periods, calibration, sensitivity to drift, and metrics related to the cost of errors. For trading applications, historical accuracy alone does not prove profitability: costs and execution constraints are needed.
- Add an LLM for a specific function. Grounds may include measurable improvement after incorporating text, a convenient natural-language interface, or reduced manual work while preserving a verifiable computational layer.
- Compare complete systems. If the LLM uses retrieval, Python, SQL, and fine-tuning, the evaluation must include the accuracy of each component, latency, cost, access control, and failure handling.
For pure numerical prediction from a financial table, classical models remain a sensible default. The strongest argument for an LLM arises not when a table can be turned into text, but when language and heterogeneous documents are genuinely part of the task. In many practical systems, the best result will come not from replacing GBDT with a language model, but from clearly separating roles: the LLM works as an interface and orchestrator, while specialized tools calculate and forecast the numbers.