Use ESM-C Online

Commercially Available ESM-C No-Code Web Server

ESM-C: Protein Language Model for Protein Representation Learning and Embeddings

ESM Cambrian (ESM-C) is a family of protein language models that learn rich representations of protein sequences, released by EvolutionaryScale (now part of Biohub). The models come in 300M, 600M and 6B parameter sizes, and the weights are released openly under the MIT license for academic and commercial use. ESM-C is a parallel model family to the ESM3 generative models, with a focus on representation learning rather than generation.

Protein language models are trained on large collections of sequences and learn patterns that reflect structure and function. Their embeddings can be used to predict properties, search for related proteins or feed downstream machine learning tasks. Better representations at lower compute cost make these models practical for more labs.

According to the release post, ESM-C is designed to deliver this quality at smaller sizes, with the 300M and 600M models offering strong results with lower memory use and faster inference than much larger predecessors.

Property

Detail

Model sizes

300M, 600M and 6B parameters

Architecture

Transformer with Pre-LN, rotary embeddings and SwiGLU activations

Training objective

Masked language modeling on protein sequences

Context length

2048 tokens

Training data

UniRef, MGnify and JGI sequences

License

MIT, for academic and commercial use

How ESM-C Is Built and Trained

ESM-C uses a standard transformer design with no biases in the linear layers or layer norms. The three models share the same training schedule of 1.5 million steps and 6.2 trillion training tokens.

Model Specifications

Parameters

Layers

Width

Heads

300M

30

960

15

600M

36

1152

18

6B

80

2560

40

Training Data and Stages

  • Sequence sources: UniRef, MGnify and Joint Genome Institute data, clustered at 70% sequence identity into 83M, 372M and 2B clusters respectively.

  • Stage 1: the first 1M steps use a context length of 512, with metagenomic data making up 64% of the data.

  • Stage 2: the final 500K steps use a context length of 2048, with metagenomic data reduced to 37.5%.

ESM-C Benchmark Results for Protein Structure Contact Prediction

The post evaluates unsupervised learning of tertiary structure by measuring contact precision (P@L) on CASP15, following an established protocol based on logistic regression over contacts.

  • ESM-C 300M: performs similarly to ESM2 650M while using less memory and running faster.

  • ESM-C 600M: rivals ESM2 3B and approaches ESM2 15B.

  • ESM-C 6B: outperforms the best ESM2 models by a wide margin.

Scaling-law results, measured on a temporally held-out PDB set, show improvement that is predictable with compute and no sign of diminishing returns up to 6B parameters.

What is Tamarind Bio?

Tamarind Bio is a no-code bioinformatics platform built to give life scientists and researchers access to powerful computational tools. Many cutting-edge machine learning models are hard to deploy and use. Tamarind provides an intuitive, web-based environment that removes the complexity of high-performance computing, software dependencies and command-line interfaces.

The platform is designed for biologists, chemists and other researchers who may not have a background in programming or cloud infrastructure but want to run models on their own data. Key features include:

  • A user-friendly graphical interface for setting up and launching experiments

  • A robust API for integration into existing research pipelines

  • An automated system for managing and scaling computational resources

Tamarind treats information and data security as a top priority, as detailed in its Trust Center and Terms of Service.

Accelerating Discovery with ESM-C on Tamarind Bio

  • Protein representations: generate embeddings of protein sequences for downstream analysis.

  • Structure-related signal: use the contact-prediction ability of the representations to probe tertiary structure.

  • Efficient inference: choose a smaller model when memory and speed matter, since the 300M and 600M models are competitive with larger ESM2 models.

  • Scaling to larger models: use the 6B model when the best representation quality is needed.

How to Use ESM-C on Tamarind Bio

  1. Log in and open the tool: sign in at tamarind.bio and open ESM-C.

  2. Provide protein sequences: enter or upload amino acid sequences, which the model reads as protein sequence input.

  3. Choose the model size: select the 300M, 600M or 6B model, if exposed.

  4. Check sequence length: keep sequences within the 2048-token context length the models were trained with.

  5. Run the job: submit it and wait for completion.

  6. Download outputs: retrieve the embeddings or other outputs the tool provides.

  7. Use downstream: feed the representations into your own prediction, clustering or search workflows.

Things to Keep in Mind

  • The reported evaluations cover only unsupervised learning of tertiary structure through contact maps.

  • The 300M and 600M models were trained beyond the compute-optimal point.

  • ESM-C is a representation model rather than a generative one; generation belongs to the ESM3 family.

Source: EvolutionaryScale, ESM Cambrian release post, https://www.evolutionaryscale.ai/blog/esm-cambrian. Code: https://github.com/evolutionaryscale/esm.

Supporting 10,000+ scientists around the world,

from leading biotechs, and global biopharma