Learning Latent Protein Languages for Autoregressive Generation

Mahdi Pourmirzaei1, Farzaneh Esmaili1, Amir Ziashahabi2, Mohammadreza Pourmirzaei3, Dong Xu1

1 University of Missouri2 University of Southern California3 Independent Researcher

Code ↗Cite this work
Tokenization and reconstruction

Amino acid sequence

Encode

PLL tokens

Decode

Reconstructed sequence

Protein backbone

Encode

SLL tokens

Decode

Reconstructed backbone

Schematic token symbols, with one token per residue.

Abstract

Autoregressive transformers are the dominant generative recipe across most tokenized modalities, yet they remain comparatively weak for protein sequence and structure generation. We study the role of target representation in this gap: amino acid tokens encode residue identities without explicit contextual semantics, while backbone coordinates require a discrete representation in our framework.

We introduce two learned latent protein languages. Protein Latent Language (PLL) maps sequences into a 4,096-state contextual alphabet built on a frozen ESM-2 encoder, retaining one token per residue. Structure Latent Language (SLL) adapts GCP-VQVAE Lite with auxiliary sequence and confidence supervision while preserving decoding to backbone coordinates. We separately pretrain autoregressive transformer language models on PLL and SLL tokens using a next-token prediction objective, yielding PLLM and SLLM. Under matched downstream sequence training, PLLM has a fitted compute-scaling exponent of 0.038 versus 0.020 for the amino acid autoregressive protein language model (PLM), a roughly 1.9× steeper slope. SLLM has a fitted exponent of 0.049 on structure-token data. In unconditional sequence generation, PLLM substantially reduces the fraction of samples below a heuristic 1.5-bit residue-composition entropy threshold by 54% relative to the amino acid model across a sweep of sampling temperatures. At moderate sampling temperatures, PLLM better matches the sequence lengths and residue-composition entropy of the UniRef50 training data than PLM does. SLL improves performance on supervised protein structure tasks. For sequence-to-structure prediction, replacing the original GCP-VQVAE Lite tokenizer with SLL reduces best validation perplexity by 34% under matched training. In backbone generation, SLLM compares favorably with other generative models on diversity and novelty. We further use SLLM for sequence-to-structure prediction, where latent-token sampling for long proteins is approximately 1000 times faster than MSA-based AlphaFold2 in our measurements, and we observe early signs that using the model's own internal token confidence for inference-time scaling can lift prediction quality beyond a single decoded sample.

Together, these results position learned latent protein languages as a promising direction for language modeling to explore, offering a practical substrate for bringing autoregressive transformer scaling and inference-time sampling to protein generation.

Method

Learn a discrete target representation, then model its tokens autoregressively.

PLL / Sequences

A contextual protein alphabet

Frozen ESM-2 representations are quantized into 4,096 states. Each residue retains one content token, preserving sequence length while giving token identities contextual information.

A two-stage VQ-RAE construction learns the codebook and then trains a transformer decoder to recover amino acid identities from the fixed codes.

SLL / Structures

Geometry with semantic supervision

SLL adapts GCP-VQVAE Lite with auxiliary residue-identity, confidence, and sequence-representation supervision. One structure token per residue remains decodable to backbone coordinates.

The auxiliary heads supervise tokenizer training. Backbone generation uses the learned codebook and structure decoder.

PLL tokenizer. Stage 1 reconstructs frozen ESM-2 representations. Stage 2 freezes the encoder and codebook and trains a reinitialized transformer decoder with a residue-classification head.

Results

Sequence generation

Less low-complexity drift

PLLM reduces the low-complexity fraction by 54% relative to the amino acid model across a sweep of sampling temperatures. At moderate temperatures, its generated sequences better match the lengths and residue-composition entropy of the UniRef50 training data.

Unfiltered generations across sampling temperatures. Residue-composition entropy is computed on decoded amino acid sequences. The UniRef50 reference marks the training distribution.
Generated lengths at sampling temperature 1.0, compared with UniRef50. PLLM more closely follows the reference distribution in this setting.

Compute scaling

A steeper fitted sequence scaling trend

Under matched downstream sequence training, the fitted loss-versus-compute exponent increases from 0.020 for the amino acid model to 0.038 for PLLM, a roughly 1.9× steeper slope.

These fits describe validation loss within each token vocabulary. They do not make raw cross-entropy values directly comparable across languages. Reported compute covers downstream language-model training.

PLM / amino acids0.020
PLLM / latent language0.038
Amino acid autoregressive model (PLM). Fitted validation loss versus downstream training compute, with exponent 0.020.
Protein Latent Language model (PLLM). The matched sequence experiment yields a fitted exponent of 0.038.
Structure-token scaling

SLLM has a fitted exponent of 0.049 in a separate structure-token experiment. It is reported independently because the sequence and structure studies use different datasets and exposures.

Structure Latent Language model (SLLM). A separate structure-token study yields a fitted exponent of 0.049.

Structure-tokenizer comparison

Better targets for sequence-to-structure modeling

SLL improves performance on supervised protein structure tasks. In a matched Prot2Token-style setup, changing the target tokenizer from GCP-VQVAE Lite to SLL lowers best validation perplexity by 34%.

Matched sequence-to-structure tokenizer comparison. Red curves show SLL, labeled GCP-VQVAE 2 Lite in the plot. Best validation perplexity is 34% lower than with GCP-VQVAE Lite. Learning-speed annotations describe progress by epoch.

Benchmarks

Protein generation results using the ProteinBench framework. Each metric captures a different aspect of the generated samples.

The length-conditioned benchmark measures folding quality, diversity, and novelty after entropy filtering. It describes a different sample population from the unfiltered drift analysis above.

Unweighted means over lengths 100, 200, 300, and 500 residues.
ModelpLDDT ↑Pairwise TM ↓Cluster ratio ↑Max TM ↓
Native63.700.51500.7900—
ProGen2‡65.260.39000.93250.6675
DPLM89.930.51000.74500.8525
ESM3‡53.240.35500.97500.6925
PLLMours59.860.25670.97000.6368
PLM†baseline65.260.24130.92500.5859

Swipe the table to see all metrics →

Our PLM and PLLM rows use sampling temperature 0.5 and 50 retained entropy-filtered samples per length, evaluated with ESMFold. Published reference rows retain ProteinBench's source protocols.

† PLM uses additional sampling when needed to fill the filtered set. ‡ ProGen2 and ESM3 use roughly 15× and 48× more protein samples than UniRef50. Highlighting follows the manuscript, with Native excluded.

Inference-time scaling

SLLM samples structure-token trajectories conditioned on E1 sequence representations. A DeepConf-style selector uses internal token confidence and token-space consensus to choose one candidate before coordinate decoding.

On held-out CAMEO-2024 targets, prediction quality improves overall as the candidate budget increases. Confidence-based selection recovers part of the gain available to an oracle that knows the reference structure.

Held-out CAMEO-2024. lDDT-CA versus candidate budget k. Higher is better. Oracle selection uses the reference structure, while DeepConf uses internal token confidence and consensus.
Held-out CAMEO-2024. RMSD versus candidate budget k. Lower is better. Oracle selection uses the reference structure, while DeepConf uses internal token confidence and consensus.
TM-score and GDT-TS
Held-out CAMEO-2024. TM-score versus candidate budget k. Higher is better. Oracle selection uses the reference structure, while DeepConf uses internal token confidence and consensus.
Held-out CAMEO-2024. GDT-TS versus candidate budget k. Higher is better. Oracle selection uses the reference structure, while DeepConf uses internal token confidence and consensus.

Oracle best@k selects using the reference structure after decoding and serves as an upper bound. DeepConf consensus@k uses the model's own token confidence before decoding. Selector settings are tuned on CASP and fixed for CAMEO evaluation.

Paper and code

Methods, detailed evaluation protocols, results by length, and limitations are available in the manuscript.

GitHub repository ↗

Citation

@misc{pourmirzaei2026latent,
  title={Learning Latent Protein Languages
         for Autoregressive Generation},
  author={Pourmirzaei, Mahdi and Esmaili, Farzaneh
          and Ziashahabi, Amir
          and Pourmirzaei, Mohammadreza and Xu, Dong},
  year={2026},
  url={https://mahdip72.github.io/latent-protein-languages.github.io/}
}

Research figureDownload image ↗