awesome-repositories.com
Blog
MCP
awesome-repositories.com

Entdecke die besten Open-Source-Repositories mit KI-gestützter Suche.

EntdeckenKuratierte SuchenOpen-Source-AlternativenSelf-hosted SoftwareBlogSitemap
ProjektMCP-ServerÜber unsRanking-MethodikPresse
RechtlichesDatenschutzAGB
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·

4 Repos

Awesome GitHub RepositoriesSemantic Segmentation Architectures

Neural network structures that combine encoders and decoders to produce pixel-wise semantic labels for images.

Distinct from Encoder-Decoder Architectures: Focuses on pixel-level semantic segmentation rather than the sequence generation found in vision-text transformer architectures

Explore 4 awesome GitHub repositories matching artificial intelligence & ml · Semantic Segmentation Architectures. Refine with filters or upvote what's useful.

Awesome Semantic Segmentation Architectures GitHub Repositories

Finde die besten Repos mit KI.Wir suchen mit KI nach den am besten passenden Repositories.
  • facebookresearch/sapiensAvatar von facebookresearch

    facebookresearch/sapiens

    5,388Auf GitHub ansehen↗

    Sapiens ist ein hochauflösendes menschliches Sichtmodell, das für hochpräzise, menschenzentrierte Computer-Vision-Aufgaben entwickelt wurde. Es fungiert als Tool-Suite zur Schätzung menschlicher Posen, Tiefe und Oberflächengeometrie. Das Projekt nutzt ein Vision-Transformer-Backbone, um mehrere Aufgaben über einen gemeinsamen Encoder auszuführen. Diese Architektur ermöglicht die gleichzeitige Vorhersage von Skelettstrukturen, Gelenkpositionen und der Entfernung zwischen einer Kamera und einer menschlichen Person. Die Funktionen des Modells decken die Segmentierung menschlicher Körperteile zur Isolierung anatomischer Regionen vom Hintergrund sowie die Vorhersage von Oberflächennormalen zur Wiederherstellung von 3D-Geometriedetails aus 2D-Bildern ab. Diese Aufgaben werden durch ein Multi-Task-Learning-Framework unterstützt, das pixelweise Regression und semantische Segmentierungsmaskierung verwendet.

    Uses neural network structures to produce pixel-wise semantic labels for isolating human subjects.

    Python
    Auf GitHub ansehen↗5,388
  • qubvel/segmentation_modelsAvatar von qubvel

    qubvel/segmentation_models

    4,917Auf GitHub ansehen↗

    This is an image segmentation framework and masking toolkit for constructing binary and multi-class neural network architectures. It serves as a deep learning encoder wrapper that integrates pre-trained convolutional neural network architectures into semantic segmentation models. The library enables the use of pre-trained backbones to isolate complex patterns and leverages transfer learning to accelerate training. It provides a collection of overlap-based loss functions and precision metrics specifically designed to evaluate and refine the accuracy of image masks. The toolkit covers the full

    Implements encoder-decoder architectures specifically for pixel-wise semantic segmentation.

    Pythondensenetefficientnetfpn
    Auf GitHub ansehen↗4,917
  • roboflow/sportsAvatar von roboflow

    roboflow/sports

    4,881Auf GitHub ansehen↗

    Roboflow Sports is a sports video analysis system that combines object detection and tracking with bird's-eye field visualization. Its core pipeline detects and tracks players, referees, and balls across video frames, then maps those tracked positions onto a radar-style overhead view of the playing field. The system goes beyond basic detection by localizing field boundaries and key landmarks such as pitch lines and corners, enabling spatial mapping of player positions relative to the field geometry. It classifies detected players by team affiliation through visual feature extraction and clust

    Classifies each pixel of video frames into field, background, or boundary categories using an encoder-decoder network.

    Pythoncomputer-visiondeep-learningdeep-neural-networks
    Auf GitHub ansehen↗4,881
  • nvlabs/segformerAvatar von NVlabs

    NVlabs/SegFormer

    3,347Auf GitHub ansehen↗

    SegFormer is a semantic segmentation framework and transformer-based model designed for pixel-level image classification. It provides a deep learning architecture that assigns class labels to pixels using a hierarchical transformer encoder and a multi-layer perceptron decoder. The framework utilizes a hierarchical transformer encoder to process multi-scale features through a pyramid of blocks and an all-MLP decoder to aggregate these features without complex attention mechanisms. It incorporates overlap patch embedding to preserve local continuity and sequential self-attention reduction to ma

    Implements a deep learning architecture that assigns class labels to pixels using a hierarchical transformer encoder and MLP decoder.

    Pythonade20kcityscapessemantic-segmentation
    Auf GitHub ansehen↗3,347
  1. Home
  2. Artificial Intelligence & ML
  3. Vision Transformers
  4. Encoder-Decoder Architectures
  5. Semantic Segmentation Architectures

Unter-Tags erkunden

  • Sports Field SegmentersEncoder-decoder networks that classify each pixel of a sports video frame into field, background, or boundary categories. **Distinct from Semantic Segmentation Architectures:** Distinct from general Semantic Segmentation Architectures: specialized for sports field pixel classification (field, background, boundary) rather than generic scene segmentation.