Tabular foundation models such as TabPFN and TabICL expect numerical matrices as input, typically formatted with shape [batch, n_samples, n_features].
Real tables, however, contain a mix of categories, short identifiers, and free-form text. Translating non-numeric columns into continuous representations requires balancing matrix width, compute budgets, and semantic fidelity.
A common point of confusion is separating how a model handles string targets () from string features (). Each follows a different set of constraints.
Recent System One models such as TypeSafe AI’s Jev introduce another option: instead of embedding an entire text field into an opaque vector, they can extract typed probabilistic judgments that become ordinary tabular features. We will get there after looking at the simpler categorical, TF-IDF, and embedding approaches.
This article outlines how tabular foundation models process string targets above their native class limits, followed by a multi-tier recipe for encoding string feature columns based on cardinality and length, and alternatives for handling long documents that exceed standard embedding context windows.
Targets vs. features: two different problems
Before encoding columns, it helps to distinguish target labels from input features:
- Target labels () produce logits. The classification head in pretrained models like
TabICLv2has a fixed output width; it emits 10 logits natively. When a dataset has 20, 40, or 100 classes, the model cannot output all class probabilities in a single pass. - Feature columns () consume attention. Feature columns do not produce logits; they enter the transformer through column and row attention. Here, the constraint is attention complexity and metric distortion rather than output head width.
Because the mechanics differ, the solutions are distinct. Targets require hierarchical prediction trees. Features require dimensionality control.
How to handle target columns with more than 10 classes
When predicting a target column with up to 10 classes, a tabular classifier uses a standard forward pass. The 10 output logits map directly to class probabilities through softmax.
When a classification target contains more than 10 classes, the model uses a hierarchical tree grouping decoder.
The algorithm operates in three steps:
- Class partitioning. The classes are recursively partitioned into coarse groups of at most 10. For example, with 40 classes, the decoder initially partitions the label space into 4 groups of 10 classes each.
- Group-level prediction. The model runs a forward pass over the training context to predict which coarse group the query row belongs to (, where ).
- Child evaluation and probability chaining. The model then evaluates the query conditioned on the selected group, predicting class probabilities within that subset ().
The final probability for class in group is the product of both stages:
This allows a foundation model trained on 10 classes to support target class counts beyond its native 10-class output without retraining.
How to encode text features for tabular foundation models
The hierarchical decoder used for targets does not solve feature encoding, because features are inputs rather than prediction outputs. Instead, feature strings are routed through four tiers based on cardinality, text length, and context budget.
| Tier | Input data type | Cardinality / length threshold | Recommended algorithm | Output width |
|---|---|---|---|---|
| Tier 1 | Low-cardinality categories | <= 10 in numeric runtimes; up to ~40 with categorical semantics | Alphabetical OrdinalEncoder | 1 column |
| Tier 2 | High-cardinality structured text | Above the chosen categorical cutoff | Character 3,4-grams + TF-IDF + Randomized SVD | 10 columns |
| Tier 3 (Small Context) | Free-form text (small context) | Sentences, paragraphs (heuristic: N < 15) | Transformer embeddings + Thin SVD | 2 to 4 columns |
| Tier 3 (Large Context) | Free-form text (large context) | Sentences, paragraphs (heuristic: N >= 15) | Transformer embeddings + PLS-1 / PLS-2 | 4 columns |
| Tier 4 | Extended documents | Multi-page text (> 512 tokens) | Jev / System One models (Score/Noul) | 10 to 20 columns |
Tier 1: Low-cardinality categories
Low-cardinality columns contain distinct discrete values such as status flags, department codes, or subscription tiers.
For these features, deterministic ordinal encoding with alphabetical sorting provides a compact representation.
Unique values: ["Pro", "Free", "Enterprise"]Sorted order: ["Enterprise", "Free", "Pro"]Integer map: {"Enterprise": 0, "Free": 1, "Pro": 2}Sorting categories alphabetically before assigning integer indices ensures consistent mappings across dataset splits, regardless of row order.
Why 10 vs. 40 as the cut-off?
Different preprocessing systems set the boundary between Tier 1 and Tier 2 at either 10 or 40 categories:
- The 40-category boundary. Tabular preprocessing libraries like
skrubtreat columns with up to 40 distinct categories as low cardinality. Below 40, categories are encoded as discrete indices or one-hot vectors. - The 10-category boundary. When the downstream runtime preserves categorical semantics, ordinal codes need not be interpreted as genuinely continuous quantities. In numeric-only runtimes, however, large arbitrary integer codes can introduce undesirable geometry (). Switching to sub-word SVD at 10 categories keeps representations smooth and avoids large discrete jumps.
Both cut-offs serve as reasonable starting heuristics. For clean nominals with little internal sub-word structure (like country codes), ordinal encoding scales reliably up to 40 categories.
Tier 2: High-cardinality structured strings
When unique values exceed 10 or 40 (such as job titles, city names, or addresses), integer encoding loses sub-word relationships. Strings like "Frontend Engineer" and "Senior Frontend Developer" share meaningful morphological overlap that integer indices discard.
A reliable recipe for this tier combines character n-gram extraction, smoothed TF-IDF weighting, and randomized SVD projection.
1. Character n-gram extraction
Instead of splitting text on full words, the tokenizer extracts character 3-grams and 4-grams with boundary padding (char_wb). For example, the string "data" expands to:
- 3-grams:
" da","dat","ata","ta " - 4-grams:
" dat","data","ata "
Character n-grams provide typo tolerance. If a value contains a minor spelling variation, the majority of its sub-word n-grams still match the vocabulary.
2. Smoothed TF-IDF weighting
Compute the inverse document frequency for every n-gram across the context rows: where is the number of rows and is the document frequency of n-gram . Normalize each row vector to unit norm.
3. Randomized SVD projection
A high-cardinality column can generate thousands of distinct n-grams. To keep matrix width compact, project the sparse TF-IDF matrix down to dimensions (with as a practical starting heuristic):
- Draw a random Gaussian test matrix .
- Form an orthonormal basis using Modified Gram-Schmidt.
- Solve the symmetric eigenvalue problem on the smaller covariance matrix (where ).
- Compute the top singular vectors: .
4. Total variance scaling
To keep projected features on the same numerical scale as other tabular columns, normalize the projection matrix by the total standard deviation across all dimensions: Dividing the projection matrix by this factor during fitting ensures that transforming new rows requires only a single matrix multiplication without additional post-processing.
Tier 3: Free-form text and document embeddings
N-gram matching reaches its limit on unstructured sentences such as support tickets or customer reviews. Phrases like "server timed out" and "database connection dropped" describe related events but share few common n-grams.
For free-form text, pre-trained transformer embedding models (such as 384-dimensional or 768-dimensional encoders) capture semantic continuity.
To prevent dense text embeddings from overwhelming the remaining tabular features, a reasonable starting heuristic is to compress the vectors down to 2 to 4 continuous dimensions.
Choosing between SVD and PLS projection
Unsupervised SVD identifies the axes of highest variance in the text embeddings, independent of target labels.
Supervised Partial Least Squares (PLS) identifies projection axes that maximize the covariance between text embeddings and target labels:
- PLS-1 handles regression and binary classification.
- PLS-2 NIPALS handles multiclass targets using one-hot indicator matrices.
The operational constraint is sample size. On small context sets, supervised PLS can overfit to training labels and degrade out-of-sample generalization.
A conservative starting heuristic:
- When training rows , use unsupervised Thin SVD.
- When training rows , use supervised PLS projection.
Setting a length threshold for zero-shot alignment
For classification tasks, computing cosine similarity between text embeddings and label descriptions can provide helpful auxiliary features.
However, instruction-tuned embedding models often prepend task prefixes to inputs. On short strings (such as email headers or short categories), this prefix can dominate the embedding vector.
As an implementation heuristic, check the average string length: When characters, skip zero-shot similarity features. Apply zero-shot alignment only to longer text blocks where document content outweighs prefix bias.
How context length limits handle long documents
When text length exceeds an embedding model’s context window, tokenizers apply silent truncation.
Feature extraction pipelines typically set truncation: true. When a text string exceeds the sequence limit, trailing tokens past the threshold are dropped prior to the model forward pass.
| Model architecture | Context window (model_max_length) | Approximate character capacity |
|---|---|---|
| Standard small encoder (e.g., MiniLM) | 512 tokens | ~1,500 to 2,000 characters |
| Larger instruction embedder (e.g., Gemma) | 2,048 tokens | ~6,000 to 8,000 characters |
Exact character capacity varies with tokenization and language.
Mean pooling operates only on the preserved tokens within the context window. The beginning of the document determines the resulting embedding.
Considerations for serialized tabular prompts
Truncation requires careful handling when using Context-Aware Semantic Embeddings (CASE), where multiple table rows are serialized into a single text prompt:
If a table contains multiple text columns and concatenates dozens of rows, the combined string can exceed 512 tokens.
Because truncation removes tokens from the end, the trailing query row may be lost. The model then evaluates historical rows without the sample being predicted.
When serializing multiple table rows into a shared prompt, keeping context row counts conservative (for example, as an operational starting heuristic) helps avoid dropping trailing query rows, or select an embedding model with a larger token window.
Why chunking doesn’t map cleanly to tabular models
When text exceeds the tokenizer context limit, the standard natural language processing solution is text chunking: split the document into smaller segments (e.g., 512 tokens each) and embed each segment independently.
In tabular foundation models, however, naive chunking creates two structural bottlenecks:
- Horizontal stacking inflates column attention cost.
If an extended document splits into 8 chunks and each chunk produces a 384-dimensional embedding, that single column expands to 3,072 dimensions. Even if each chunk is compressed down to 4 components via SVD, that field still occupies tabular columns. Tabular transformers pay a computational and statistical cost as feature width grows. Although
TabICLv2was pretrained on tables with up to 100 columns and can generalize substantially beyond that, turning one text field into dozens or hundreds of artificial columns dilutes cross-column attention and increases inference latency. - Vertical chunk pooling washes out localized facts. Averaging chunk embeddings into a single vector preserves column width, but mean pooling across multiple paragraphs flattens critical localized signals, such as an explicit penalty clause in paragraph 8 or a specific warranty condition.
Tier 4: Jev and System One models as semantic feature extractors
When text length exceeds standard tokenizer limits or documents contain complex multi-page narratives, an effective alternative to abstract vector embeddings is structured semantic featurization, using models like jev-1.12 via TypeSafe’s SystemOne interface, as detailed in TypeSafe’s Autoresearch Feature Discovery cookbook.
Rather than compressing text into an opaque numerical vector or fragmenting it into chunks, an evaluator model reads the entire document within its native context window and evaluates structured questions against defined rubrics.
Concrete implementation: Score and Noul questions
In TypeSafe’s SystemOne API, text features are evaluated through two primary question primitives:
Score(instructions, criteria): Evaluates a multi-point intensity rubric (such as a 5-point scale). Instead of returning free-form text, the evaluator outputs a probability distribution across discrete levels .Noul(instructions, criteria): Evaluates a boolean presence or condition check, outputting a probability scalar .
Here is a concrete Python example querying jev-1.12 to extract structured features from a long document (such as a customer escalation or contract review):
from typesafe_sdk import TypeSafeClient, Score, Noul, NoulCriteria
client = TypeSafeClient(api_key="your-api-key")
# Multi-level rubric for intensity featuresSEVERITY_LEVELS = [ "No issue reported or tone is positive", "Minor friction mentioned in passing", "Moderate issue with product operation", "Severe disruption affecting active business", "Critical failure with explicit threat of churn or legal action",]
# Criteria for binary presence checksCANCELLATION_CRITERIA = NoulCriteria( true="The customer explicitly requests cancellation, refund, or termination", false="No explicit cancellation or refund request is made",)
# Long document (can span thousands of tokens without chunking)document_text = """We have been experiencing repeated connection drops with your API since yesterday...[Several pages of error logs, email exchanges, and timeline details]...Please issue a full refund for this month's invoice and cancel our enterprise tier."""
response = client.system_one( model="jev-1.12", state=document_text, questions={ "severity": Score( instructions="How severe is the business disruption described by the customer?", criteria=SEVERITY_LEVELS, ), "requested_cancellation": Noul( instructions="Does this communication request a contract cancellation or refund?", criteria=CANCELLATION_CRITERIA, ), },)
# Convert answers into scalar tabular columnsseverity_probs = [response.answers["severity"].probabilities.get(i, 0.0) for i in range(5)]expected_severity = sum(i * p for i, p in enumerate(severity_probs))severity_spread = (sum(((i - expected_severity) ** 2) * p for i, p in enumerate(severity_probs))) ** 0.5
prob_cancellation = response.answers["requested_cancellation"].noul
# Resulting tabular columns for TabICLv2:# [expected_severity, severity_spread, prob_cancellation] -> 3 clean numerical floatsFrom question probabilities to tabular columns
The response yields probabilities that map directly to continuous columns:
- Expected rubric score:
- Score variance / uncertainty:
- Binary probability:
A set of 10 to 15 targeted questions produces 10 to 25 continuous columns. This avoids column explosion and provides compact, semantically meaningful numerical inputs for tabular foundation models.
This changes the role of the language model. Instead of asking it to make the final prediction, we use it as a semantic encoder: it maps unstructured evidence into a small set of probabilistic, interpretable variables. The tabular foundation model still learns how those variables interact with the rest of the table.
Jev vs. embeddings for tabular features
Comparing dense text embeddings to System One models like Jev reveals distinct architectural trade-offs for tabular foundations:
- Dense text embeddings: Provide a generic, unsupervised latent representation. They are compact, fast, and can run locally in-browser on WebGPU, but they are task-agnostic, uninterpretable, and awkward to chunk when documents exceed token windows.
- Jev / System One models: Act as task-conditioned semantic encoders. They evaluate long documents within their native context window to extract structured probabilistic variables, yielding bounded tabular columns that models like TabPFN and TabICL can directly consume.
| Metric | Chunked Dense Embeddings | Jev / System One Features |
|---|---|---|
| Column width consumption | 16 to 64+ columns (high attention cost) | Exactly 1 to 2 columns per question (typically 10 to 20 total) |
| Document length handling | Truncated or split across chunks | Evaluated across full length in the model context window |
| Target alignment | Unsupervised (captures syntax, topic, and boilerplate) | Target-conditioned (questions focus on factors predictive of ) |
| Interpretability | Opaque linear projections (dim_0, dim_1) | Domain-specific continuous metrics (refund_prob, urgency_score) |
By translating long-form text that fits within the evaluator’s context window into 10 to 20 dense, continuous scalar features bounded between and , tabular foundation models receive high-density signal without exceeding column attention budgets.
Operational trade-offs
While question-based feature extraction preserves tabular column budgets, it introduces distinct operational considerations:
- Preventing target leakage during question discovery. Because question discovery can use target labels to identify high-signal questions, question generation and selection must be fitted strictly on the training split. Evaluating or selecting questions against the complete dataset introduces target leakage.
- Automated question discovery requires offline compute. Discovering which questions best predict the target (the autoresearch loop) involves proposing candidate questions, scoring training rows, and running feature importance tests. This is well-suited for dataset onboarding and batch pipelines, but too computationally intensive for instantaneous ad-hoc spreadsheet drag-and-drops.
- Runtime execution context. Small embedding models can run interactively in-browser on modern WebGPU hardware, whereas large evaluator models require dedicated GPU infrastructure (such as Cloud Run or server-side inference) to process batches of multi-page documents.
Summary of encoding guidelines
- Separate targets from features. Use hierarchical tree grouping for target variables exceeding 10 classes; reserve dimensionality reduction for input features.
- Keep low-cardinality features compact. Deterministic alphabetical ordinal encoding preserves matrix width for low-cardinality categories, scaling up to 10 or 40 values depending on whether the runtime preserves categorical semantics.
- Use character n-grams and SVD for structured text. Sub-word TF-IDF with randomized SVD captures typo tolerance and morphological overlap, with 10 dimensions serving as a practical default.
- Use dense embeddings for moderate text. For sentences and short paragraphs within 512 tokens, compress embeddings to 2 to 4 features, fall back to SVD on small sample sizes (heuristic: ), and monitor context limits during prompt serialization.
- Use Jev / System One models for long documents. When documents exceed standard embedding windows, avoid naive horizontal chunk stacking that inflates column attention cost. Extracting scalar probabilities from targeted questions preserves document-wide context while keeping tabular column counts strictly bounded.
How Carla handles this in your browser
Wiring up custom text tokenizers, tuning SVD components, and managing hierarchical class decoders takes engineering time.
Carla handles this routing automatically from the browser. Tabular inference with TabPFN and TabICLv2 runs locally on WebGPU, while long-document semantic extraction with Jev can connect to a dedicated evaluator service when needed. When an external evaluator is used, only the text required for that step is sent to the configured service.
You can try TabICLv2 directly on your spreadsheets with the Carla Chrome extension or review our interactive benchmark suite.
