awesome-repositories.com
Blog
MCP
awesome-repositories.com

Discover the best open-source repositories with AI-powered search.

ExploreCurated searchesOpen-source alternativesSelf-hosted softwareBlogSitemap
ProjectMCP serverAboutHow we rankPress
LegalPrivacyTerms
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
ArcInstitute avatar

ArcInstitute/evo2

0
View on GitHub↗
3,951 stars·505 forks·Jupyter Notebook·Apache-2.0·16 views

Evo2

evo2 is a genomic large language model and foundation model designed to predict, generate, and analyze genetic information across different species. It functions as a nucleotide sequence modeler and a DNA sequence generator, using transformer-based sequence modeling to process genomic data.

The system provides capabilities for synthetic DNA generation, creating new genetic sequences based on biological prompts or species-specific tags. It also performs nucleotide likelihood prediction to score genomic variants and analyze biological properties within DNA sequences.

The model supports genomic sequence analysis through the extraction of high-dimensional representations from intermediate layers. These embeddings enable specialized classification and downstream analysis of genetic data.

Features

  • Genomic Sequence Modeling - Uses transformer-based self-attention mechanisms to predict nucleotide likelihoods and capture long-range dependencies in genomic data.
  • Synthetic DNA Generation - Creates new genetic sequences based on specific prompts or species tags to fill gaps in genomic data.
  • Genomic Foundation Models - Serves as a pre-trained biological model providing high-dimensional sequence embeddings for downstream genomic analysis.
  • Genomic LLMs - Implements a large language model trained on DNA sequences to predict, generate, and analyze genetic information.
  • Genomic Cross-Domain Pretraining - Learns general biological patterns from diverse species data to allow a single model to generalize across all domains of life.
  • Genomic Sequences - Produces new DNA sequences by iteratively predicting the next nucleotide based on biological tokens and species tags.
  • DNA Sequence Generators - Provides a generative model that produces synthetic genetic sequences based on biological prompts or species tags.
  • Genomic - Produces new genetic sequences based on prompts or species tags to complete missing sequence information.
  • Genomic Data Analysis - Extracts high-dimensional embeddings from genetic data to perform specialized biological classification and sequence analysis.
  • Genomic Sequence Interpreters - Predicts nucleotide likelihoods across a sequence to score genomic variants or analyze biological properties.
  • Genome Modeling and Design - Uses machine learning to predict and create DNA sequences for research and biological engineering across species.
  • Nucleotide Likelihood Prediction - Predicts the probability of specific nucleotides in a sequence to score genomic variants and biological properties.
  • Layer Extractions - Captures high-dimensional representations from intermediate model layers for specialized downstream biological analysis.
  • Genomic Sequence Embeddings - Captures high-dimensional representations from intermediate model layers for specialized analysis of sequence data.
  • Genomic Tokenization - Converts raw nucleotide sequences into discrete tokens that the model processes as a structured vocabulary.

Star history

Star history chart for arcinstitute/evo2Star history chart for arcinstitute/evo2

How this analysis was created: This summary and feature list are AI-generated from collected project material and can contain mistakes. Stars, license and language are imported from GitHub. Inclusion does not mean that we have tested or audited this project. Check the source documentation for any feature you depend on. Learn more on our About page.

AI search

Explore more awesome repositories

Describe what you need in plain English — the AI ranks thousands of curated open-source projects by relevance.

Start searching with AI

Projects sharing features with Evo2

These projects share indexed features with Evo2. Shared tags can include platform or build tooling; verify the primary use case before treating a result as a replacement.
  • biopython/biopythonbiopython avatar

    biopython/biopython

    5,078View on GitHub↗

    Biopython is a bioinformatics library for Python providing tools to parse, manipulate, and analyze biological sequences, molecular structures, and phylogenetic trees. It serves as a biological sequence parser for genomic and proteomic data across multiple industry-standard file formats and acts as an interface for querying biological data and citations from NCBI Entrez repositories. The project distinguishes itself through specialized toolkits for protein structure analysis and phylogenetic tree construction. It includes a protein structure analyzer for processing PDB and mmCIF files to calcu

    Pythonbioinformaticsbiopythondna
    View on GitHub↗5,078
  • k-dense-ai/claude-scientific-skillsK-Dense-AI avatar

    K-Dense-AI/claude-scientific-skills

    8,907View on GitHub↗

    This project is a scientific agent framework and workflow orchestrator designed to extend large language models with specialized tools for genomic, chemical, and biological research. It provides a system for planning research hypotheses and executing automated workflows by integrating scientific databases with dynamic code execution. The framework includes a cheminformatics modeling suite for predicting molecular bioactivity and performing virtual screening, alongside a bioinformatics analysis toolkit for processing genomic sequences and single-cell data. It also features an academic document

    Pythonai-scientistbioinformaticschemoinformatics
    View on GitHub↗8,907
  • google-research/google-researchgoogle-research avatar

    google-research/google-research

    38,139View on GitHub↗

    This repository serves as a comprehensive research platform and toolkit for advancing machine learning, quantum computing, and large-scale scientific data analysis. It provides foundational frameworks for developing complex algorithmic systems, offering the necessary infrastructure for distributed training, computational graph execution, and high-performance model development. The project distinguishes itself by integrating specialized research domains with robust, privacy-preserving methodologies. It supports diverse scientific discovery through tools for quantum simulation, physics-informed

    Jupyter Notebookaimachine-learningresearch
    View on GitHub↗38,139
  • hail-is/hailhail-is avatar

    hail-is/hail

    1,064View on GitHub↗

    Cloud-native genomic dataframes and batch computing

    Python
    View on GitHub↗1,064
Compare all 7 related projects→

Frequently asked questions

What does arcinstitute/evo2 do?

evo2 is a genomic large language model and foundation model designed to predict, generate, and analyze genetic information across different species. It functions as a nucleotide sequence modeler and a DNA sequence generator, using transformer-based sequence modeling to process genomic data.

What are the main features of arcinstitute/evo2?

The main features of arcinstitute/evo2 are: Genomic Sequence Modeling, Synthetic DNA Generation, Genomic Foundation Models, Genomic LLMs, Genomic Cross-Domain Pretraining, Genomic Sequences, DNA Sequence Generators, Genomic.

Which projects share features with arcinstitute/evo2?

Projects with overlapping indexed features include: biopython/biopython — Biopython is a bioinformatics library for Python providing tools to parse, manipulate, and analyze biological… k-dense-ai/claude-scientific-skills — This project is a scientific agent framework and workflow orchestrator designed to extend large language models with… google-research/google-research — This repository serves as a comprehensive research platform and toolkit for advancing machine learning, quantum… hail-is/hail — Cloud-native genomic dataframes and batch computing. dnanexus-rnd/glnexus — Scalable gVCF merging and joint variant calling for population sequencing projects. tingsongyu/pytorch-tutorial-2nd — This project is a comprehensive instructional resource and course for building neural networks using PyTorch. It…