awesome-repositories.com
المدونة
awesome-repositories.com

اكتشف أفضل مستودعات المصادر المفتوحة باستخدام بحث مدعوم بالذكاء الاصطناعي.

استكشفعمليات بحث منسقةبدائل مفتوحة المصدربرمجيات ذاتية الاستضافةالمدونةخريطة الموقع
المشروعحولكيفية ترتيب النتائجالصحافةخادم MCP
قانونيالخصوصيةالشروط
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
Back to sdv-dev/sdv

Open-source alternatives to SDV

30 open-source projects similar to sdv-dev/sdv, ranked by how many features they have in common. Compare stars, activity and what each one does to find the best SDV alternative.

  • ahupp/python-magicالصورة الرمزية لـ ahupp

    ahupp/python-magic

    2,886عرض على GitHub↗

    python-magic is a C-binding wrapper that provides a Python interface for the libmagic system library. It functions as a file signature analyzer and MIME type detector, identifying file formats by comparing header bytes against a database of known binary signatures. The library enables the identification of file types from both file paths and raw data buffers. It supports custom file signature matching through the injection of user-provided magic databases, allowing for the detection of specialized or proprietary formats. The project covers binary data analysis and MIME type mapping to transl

    Python
    عرض على GitHub↗2,886
  • argilla-io/argillaالصورة الرمزية لـ argilla-io

    argilla-io/argilla

    5,015عرض على GitHub↗

    Argilla is a collaborative AI feedback tool and data curation management system. It serves as a human-in-the-loop dataset platform designed to coordinate workforce annotators and domain experts in labeling, rating, and refining data samples for machine learning projects. The platform focuses on large language model dataset curation and reinforcement learning from human feedback workflows. It provides a shared workspace for integrating human expertise into AI development to validate model outputs and correct data errors. The system manages the end-to-end machine learning data pipeline, includ

    Python
    عرض على GitHub↗5,015
  • borb-pdf/borbB

    borb-pdf/borb

    0عرض على GitHub↗
    عرض على GitHub↗0
  • camelot-dev/camelotالصورة الرمزية لـ camelot-dev

    camelot-dev/camelot

    3,764عرض على GitHub↗

    Camelot is a Python library and processing engine designed to extract tabular data from PDF documents. It converts unstructured tables into machine-readable formats such as CSV, JSON, and Excel. The project provides specialized toolsets for different document types, using line detection for ruled tables and whitespace analysis for borderless tables. It includes an optical character recognition system to recover structured data from image-based scanned PDFs that lack a digital text layer. The library handles complex document layouts, including encrypted files, rotated pages, and tables that s

    Python
    عرض على GitHub↗3,764

بحث بالذكاء الاصطناعي

استكشف المزيد من المستودعات الرائعة

صف ما تحتاجه بلغة بسيطة — وسيقوم الذكاء الاصطناعي بترتيب آلاف المشاريع مفتوحة المصدر المنسقة حسب الصلة.

Find more with AI search
  • camelot-dev/excaliburالصورة الرمزية لـ camelot-dev

    camelot-dev/excalibur

    1,790عرض على GitHub↗
    Pythonextractfor-humanspdf
    عرض على GitHub↗1,790
  • cleanlab/cleanlabالصورة الرمزية لـ cleanlab

    cleanlab/cleanlab

    11,513عرض على GitHub↗

    Cleanlab is a data-centric AI library and toolkit designed to improve machine learning model performance by detecting label errors and increasing overall dataset quality. It implements a confident learning framework that iteratively refines label noise estimates by comparing model predictions with estimated label probabilities to identify mislabeled examples. The project provides specialized utilities for active learning optimization, allowing for the selection of the most impactful examples for labeling or re-labeling. It also includes an outlier detection tool to identify atypical data poin

    Pythonactive-learningannotationanomaly-detection
    عرض على GitHub↗11,513
  • code-kern-ai/refineryالصورة الرمزية لـ code-kern-ai

    code-kern-ai/refinery

    1,470عرض على GitHub↗

    The data scientist's open-source choice to scale, assess and maintain natural language data. Treat training data like a software artifact.

    Python
    عرض على GitHub↗1,470
  • cvat-ai/cvatالصورة الرمزية لـ cvat-ai

    cvat-ai/cvat

    15,317عرض على GitHub↗

    CVAT is an open-source, web-based platform designed for annotating images, videos, and 3D point clouds to create high-quality training datasets for machine learning. It functions as a containerized server that orchestrates the entire lifecycle of computer vision data, from initial task creation and manual labeling to quality assurance and final dataset export. The platform distinguishes itself through deep integration with machine learning models, allowing users to deploy custom AI models as serverless functions for automated object detection, tracking, and skeleton annotation. It supports co

    Pythonannotationannotation-toolannotations
    عرض على GitHub↗15,317
  • deanmalmgren/textractالصورة الرمزية لـ deanmalmgren

    deanmalmgren/textract

    4,623عرض على GitHub↗

    Textract is a multi-format text extraction tool and parser. It provides a unified interface to extract plain text from a variety of sources, including documents, images, and audio files. The system functions as a document content parser for PDFs and spreadsheets, an image text extractor using optical character recognition, and a speech-to-text transcriber for audio recordings.

    HTML
    عرض على GitHub↗4,623
  • diyago/tabular-data-generationالصورة الرمزية لـ Diyago

    Diyago/Tabular-data-generation

    570عرض على GitHub↗

    We well know GANs for success in the realistic image generation. However, they can be applied in tabular data generation. We will review and examine some recent papers about tabular GANs in action.

    Pythonadversarial-filteringdeep-learningfeature-engineering
    عرض على GitHub↗570
  • doccano/doccanoالصورة الرمزية لـ doccano

    doccano/doccano

    10,674عرض على GitHub↗

    Doccano is a collaborative data labeling platform and machine learning dataset management system. It provides a web-based interface for teams to import raw text, mark datasets, and export structured annotations for model training. The project specifically supports text annotation for classification and named entity recognition tasks. It enables teams to coordinate multiple users on a single project to maintain consistent labeling guidelines and increase the speed of dataset creation. The system includes tools for data management and team coordination, providing the ability to import raw data

    Python
    عرض على GitHub↗10,674
  • frictionlessdata/tabulator-pyF

    frictionlessdata/tabulator-py

    0عرض على GitHub↗
    عرض على GitHub↗0
  • gretelai/gretel-syntheticsالصورة الرمزية لـ gretelai

    gretelai/gretel-synthetics

    679عرض على GitHub↗

    Synthetic data generators for structured and unstructured text, featuring differentially private learning.

    Python
    عرض على GitHub↗679
  • hitachi-automotive-and-industry-lab/semantic-segmentation-editorالصورة الرمزية لـ Hitachi-Automotive-And-Industry-Lab

    Hitachi-Automotive-And-Industry-Lab/semantic-segmentation-editor

    1,961عرض على GitHub↗

    Web labeling tool for bitmap images and point clouds

    JavaScript
    عرض على GitHub↗1,961
  • huggingface/datasetsالصورة الرمزية لـ huggingface

    huggingface/datasets

    21,643عرض على GitHub↗

    Datasets is a library designed for the management, processing, and sharing of large-scale data collections for machine learning workflows. It functions as both a data processing framework and a versioning platform, providing tools to organize, filter, and transform massive datasets while ensuring reproducibility across research and development teams. The library distinguishes itself by enabling the handling of datasets that exceed available system memory. It utilizes memory-mapped file access, disk-based caching, and lazy iterative streaming to maintain performance when working with large-sca

    Pythonaiartificial-intelligencecomputer-vision
    عرض على GitHub↗21,643
  • humansignal/label-studioالصورة الرمزية لـ HumanSignal

    HumanSignal/label-studio

    27,619عرض على GitHub↗

    Label Studio is a multi-modal data annotation platform designed to create and manage high-quality training datasets for machine learning. It functions as a self-hosted, containerized environment that supports secure, private deployments, including air-gapped configurations. The platform provides a centralized workspace for labeling diverse media types, such as images, text, audio, and time-series data, to support supervised and reinforcement learning workflows. The platform distinguishes itself through deep integration with machine learning backends, enabling active learning loops, automated

    TypeScriptannotationannotation-toolannotations
    عرض على GitHub↗27,619
  • intake/intakeالصورة الرمزية لـ intake

    intake/intake

    1,080عرض على GitHub↗

    Intake is a lightweight package for finding, investigating, loading and disseminating data.

    Python
    عرض على GitHub↗1,080
  • jazzband/tablibالصورة الرمزية لـ jazzband

    jazzband/tablib

    4,754عرض على GitHub↗

    Tablib is a Python library designed for importing, exporting, and manipulating tabular datasets. It functions as a multi-format data converter and manager, allowing users to move information between different file standards. The library supports data transformation across CSV, JSON, YAML, and Excel formats. It provides a programmatic interface to manage these datasets by adding rows, filtering columns, and segregating records. The system uses a common internal representation and adapter-based mapping to normalize diverse input sources. This allows for consistent reading and writing routines

    Python
    عرض على GitHub↗4,754
  • joke2k/fakerالصورة الرمزية لـ joke2k

    joke2k/faker

    19,278عرض على GitHub↗

    Faker is a Python library designed to generate realistic synthetic data for software testing, database prototyping, and privacy-preserving anonymization. It provides a comprehensive suite of tools to create diverse information types, including personal identities, financial records, geographic locations, and technical system metadata, allowing developers to populate environments with mock data that mimics real-world structures. The library is built on a modular provider architecture that supports dynamic method dispatch, enabling users to extend functionality by registering custom data genera

    Pythondatasetfakefake-data
    عرض على GitHub↗19,278
  • jsbroks/coco-annotatorالصورة الرمزية لـ jsbroks

    jsbroks/coco-annotator

    2,276عرض على GitHub↗

    :pencil2: Web-based image segmentation tool for object detection, localization, and keypoints

    Vue
    عرض على GitHub↗2,276
  • lightly-ai/lightly-studioالصورة الرمزية لـ lightly-ai

    lightly-ai/lightly-studio

    848عرض على GitHub↗

    Curate, Annotate, and Manage Your Data in LightlyStudio.

    Python
    عرض على GitHub↗848
  • martinblech/xmltodictالصورة الرمزية لـ martinblech

    martinblech/xmltodict

    5,741عرض على GitHub↗

    xmltodict is a Python library that provides bidirectional serialization between XML documents and dictionaries. It functions as a parser that converts marked-up input into key-value pairs and a serialization utility that transforms dictionaries back into structured XML documents. The project includes an incremental stream processor that uses depth-based callbacks to handle large XML files while maintaining constant memory usage. It features a namespace manager for mapping prefixes and declarations, as well as a security sanitizer that blocks external entity expansion and validates element nam

    Python
    عرض على GitHub↗5,741
  • merantix-momentum/squirrel-coreM

    merantix-momentum/squirrel-core

    0عرض على GitHub↗
    عرض على GitHub↗0
  • nv-tlabs/vipeالصورة الرمزية لـ nv-tlabs

    nv-tlabs/vipe

    1,990عرض على GitHub↗

    ViPE: Video Pose Engine for Geometric 3D Perception

    Python
    عرض على GitHub↗1,990
  • nvidia/nemo-curatorالصورة الرمزية لـ NVIDIA

    NVIDIA/NeMo-Curator

    1,620عرض على GitHub↗

    Scalable data pre processing and curation toolkit for LLMs

    Python
    عرض على GitHub↗1,620
  • pydata/pandas-datareaderالصورة الرمزية لـ pydata

    pydata/pandas-datareader

    3,217عرض على GitHub↗

    Extract data from a wide range of Internet sources into a pandas DataFrame.

    Pythondatadata-analysisdataset
    عرض على GitHub↗3,217
  • pyexcel/pyexcel-xlsxP

    pyexcel/pyexcel-xlsx

    0عرض على GitHub↗
    عرض على GitHub↗0
  • python-excel/xlrdالصورة الرمزية لـ python-excel

    python-excel/xlrd

    2,205عرض على GitHub↗

    Please use openpyxl where you can...

    Python
    عرض على GitHub↗2,205
  • rom1504/img2datasetالصورة الرمزية لـ rom1504

    rom1504/img2dataset

    4,423عرض على GitHub↗

    img2dataset is a high-performance image dataset pipeline and preprocessing tool designed to download and process millions of images from URLs for machine learning training. It functions as a distributed image downloader and cloud storage data exporter, moving large visual datasets from web sources directly into structured formats. The system prioritizes high-throughput data acquisition by distributing workloads across multiple CPU cores and machines. It integrates directly with remote cloud storage buckets and employs a manifest-based tracking system to resume interrupted downloads without re

    Pythonbig-datadatasetdeep-learning
    عرض على GitHub↗4,423
  • simonw/csvs-to-sqliteالصورة الرمزية لـ simonw

    simonw/csvs-to-sqlite

    932عرض على GitHub↗

    Convert CSV files into a SQLite database. Browse and publish that SQLite database with Datasette.

    Python
    عرض على GitHub↗932