awesome-repositories.com
المدونة
awesome-repositories.com

اكتشف أفضل مستودعات المصادر المفتوحة باستخدام بحث مدعوم بالذكاء الاصطناعي.

استكشفعمليات بحث منسقةبدائل مفتوحة المصدربرمجيات ذاتية الاستضافةالمدونةخريطة الموقع
المشروعحولكيفية ترتيب النتائجالصحافةخادم MCP
قانونيالخصوصيةالشروط
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
huggingface avatar

huggingface/text-embeddings-inference

0
View on GitHub↗
4,871 نجوم·399 تفرعات·Rust·Apache-2.0·8 مشاهداتhuggingface.co/docs/text-embeddings-inference/quick_tour↗

Text Embeddings Inference

Text Embeddings Inference هو خادم استدلال عالي الأداء مصمم لاستضافة نماذج تضمين النصوص وتصنيف التسلسلات كنقاط نهاية API قابلة للتوسع. يوفر واجهة برمجة تطبيقات لتضمين المتجهات لتحويل النص إلى تمثيلات كثيفة وخادم إعادة ترتيب (reranking) عبر المشفرات لتسجيل مدى صلة تسلسلات المستندات مقابل استعلام.

يتميز المشروع بمحرك استدلال مسرع بواسطة GPU يستخدم التجميع الديناميكي والنواة المتخصصة لزيادة الإنتاجية. يوفر واجهة ثنائية عالية الأداء عبر gRPC كبديل لـ HTTP القياسي لتقليل زمن انتقال الشبكة وتكاليف التسلسل.

يغطي النظام مجموعة واسعة من القدرات، بما في ذلك ترتيب تشابه المستندات، وإعادة ترتيب النصوص متعددة اللغات، وتصنيف التسلسلات للتنبؤ بالفئات أو المشاعر. يدعم بيئات نشر متنوعة، تتراوح من حاويات التوسع التلقائي بدون خادم إلى التثبيتات المعزولة عن الإنترنت.

يتوفر تسريع الأجهزة لوحدات معالجة الرسومات NVIDIA و AMD و Apple Metal.

Features

  • High Throughput Inference - Optimizes hardware utilization with dynamic batching and GPU acceleration to process large volumes of text requests.
  • Embedding Servers - Acts as a high-performance network-accessible server for generating high-dimensional vector representations.
  • Dynamic Batching Engines - Groups multiple individual requests into a single GPU operation to maximize hardware throughput and reduce compute overhead.
  • Document Rerankers - Provides a dedicated reranking server to score the relevance of documents relative to a query across multiple languages.
  • GPU-Accelerated Inference - Utilizes a compute-optimized runtime with dynamic batching and specialized kernels to maximize GPU throughput.
  • Throughput Optimizers - Implements high-throughput inference using dynamic batching and optimized transformer kernels to maximize GPU utilization.
  • Cross-Encoder Rerankers - Implements a high-precision reranking server using joint cross-attention scoring for query-document pairs.
  • Retrieval Re-ranking - Uses cross-encoder models to score the relevance of candidate texts against a query to improve search accuracy.
  • Specialized Mathematical Kernels - Uses specialized low-level mathematical implementations to accelerate the core attention and linear layers of embedding models.
  • Text Classification - Serves models that categorize and label text inputs for predicting sentiment or emotions.
  • Text Embedding Generators - Transforms text inputs into dense vector representations using instruction-based queries.
  • Text Embeddings - Provides a scalable server for generating dense vector representations from text for AI applications.
  • Vector Embeddings - Provides an API that converts text into dense vector representations for semantic search.
  • Cross-Encoder Scoring - Uses cross-encoder models to calculate similarity scores for query-passage pairs to refine search results.
  • ML Model Hosting - Supports loading and hosting model weights from remote repositories or local directories.
  • AMD Hardware Acceleration - Enables embedding and classification models to run on AMD hardware using the ROCm compute platform.
  • Apple Hardware Acceleration - Provides hardware acceleration for embedding and classification models locally on macOS using Apple Silicon.
  • RAG Pipeline Integrations - Generates high-performance embeddings and similarity scores to power retrieval-augmented generation pipelines.
  • Hardware Dispatchers - Directs model execution to optimized compute paths for NVIDIA GPUs, AMD GPUs, or Apple Metal based on the host.
  • Long Context Processing - Supports the processing of large input sequences consisting of several thousand tokens for extended document context.
  • Memory-Mapped Weight Loaders - Maps model weights directly from disk into memory to enable rapid startup and efficient resource sharing across processes.
  • Similarity Query Engines - Uses cross-encoder models to rank the semantic closeness of documents relative to a query.
  • Batch Input Processing - Handles multiple text inputs in a single request to increase total inference throughput.
  • Model Endpoint Deployment - Creates hosted environments with specific hardware accelerators and runtime configurations for model inference.
  • Air-Gapped Deployments - Supports running models in isolated networks by mounting pre-downloaded weights as volumes.
  • gRPC Interfaces - Implements a high-performance binary interface via gRPC to reduce serialization latency and network overhead.
  • Text Embedding gRPC Services - Provides a high-performance gRPC binary interface to reduce network latency compared to standard HTTP.
  • Concurrent Request Limits - Limits the number of simultaneous requests per instance to prevent timeouts and manage back pressure.
  • Inference and Serving - Optimized inference for text embeddings in Rust.
  • Model Serving Engines - Dedicated inference server for text-embedding models.

سجل النجوم

مخطط تاريخ النجوم لـ huggingface/text-embeddings-inferenceمخطط تاريخ النجوم لـ huggingface/text-embeddings-inference

بحث بالذكاء الاصطناعي

استكشف المزيد من المستودعات الرائعة

صف ما تحتاجه بلغة بسيطة — وسيقوم الذكاء الاصطناعي بترتيب آلاف المشاريع مفتوحة المصدر المنسقة حسب الصلة.

Start searching with AI

الأسئلة الشائعة

ما هي وظيفة huggingface/text-embeddings-inference؟

Text Embeddings Inference هو خادم استدلال عالي الأداء مصمم لاستضافة نماذج تضمين النصوص وتصنيف التسلسلات كنقاط نهاية API قابلة للتوسع. يوفر واجهة برمجة تطبيقات لتضمين المتجهات لتحويل النص إلى تمثيلات كثيفة وخادم إعادة ترتيب (reranking) عبر المشفرات لتسجيل مدى صلة تسلسلات المستندات مقابل استعلام.

ما هي الميزات الرئيسية لـ huggingface/text-embeddings-inference؟

الميزات الرئيسية لـ huggingface/text-embeddings-inference هي: High Throughput Inference, Embedding Servers, Dynamic Batching Engines, Document Rerankers, GPU-Accelerated Inference, Throughput Optimizers, Cross-Encoder Rerankers, Retrieval Re-ranking.

ما هي البدائل مفتوحة المصدر لـ huggingface/text-embeddings-inference؟

تشمل البدائل مفتوحة المصدر لـ huggingface/text-embeddings-inference: pytorch/serve — This project is a PyTorch model serving framework designed to deploy and scale machine learning models in production… kreuzberg-dev/kreuzberg — Kreuzberg is a document extraction engine that converts PDFs, Office files, images, and over 90 other formats into… johnsnowlabs/spark-nlp — Spark NLP is a toolkit for scalable text analysis and machine learning built on the Apache Spark distributed computing… hanxiao/bert-as-service — This project is a high-performance BERT embedding service and inference server designed to map text sequences into… ollama/ollama-js — ollama-js is a JavaScript client library and API wrapper that provides a programmatic interface for interacting with… openvinotoolkit/openvino — OpenVINO is an AI inference engine and model serving platform designed to execute optimized deep learning models…

بدائل مفتوحة المصدر لـ Text Embeddings Inference

مشاريع مفتوحة المصدر مشابهة، مرتبة حسب عدد الميزات المشتركة مع Text Embeddings Inference.
  • pytorch/serveالصورة الرمزية لـ pytorch

    pytorch/serve

    4,354عرض على GitHub↗

    This project is a PyTorch model serving framework designed to deploy and scale machine learning models in production via scalable network endpoints. It functions as a high-performance inference server, optimizer, and model lifecycle manager that handles model loading, request batching, and hardware acceleration. The system distinguishes itself through advanced orchestration and optimization capabilities, such as chaining multiple models into sequential workflows using execution graphs and employing dynamic batching to improve throughput and latency. It provides specialized support for generat

    Java
    عرض على GitHub↗4,354
  • kreuzberg-dev/kreuzbergالصورة الرمزية لـ kreuzberg-dev

    kreuzberg-dev/kreuzberg

    8,527عرض على GitHub↗

    Kreuzberg is a document extraction engine that converts PDFs, Office files, images, and over 90 other formats into clean, structured text and metadata. It is built around a compiled Rust core that can be used as a native library, a command-line tool, a REST API server, or a WebAssembly module for browser-based processing. The system is designed to run entirely on self-hosted infrastructure, with no data leaving the user's environment. What distinguishes Kreuzberg is its breadth of integration surfaces and its pipeline architecture. It exposes extraction capabilities through native bindings fo

    Rustdocument-intelligenceelixirffi
    عرض على GitHub↗8,527
  • johnsnowlabs/spark-nlpالصورة الرمزية لـ JohnSnowLabs

    JohnSnowLabs/spark-nlp

    4,135عرض على GitHub↗

    Spark NLP is a toolkit for scalable text analysis and machine learning built on the Apache Spark distributed computing framework. It provides a multimodal machine learning framework and a distributed pipeline system for sequencing annotators to process large-scale linguistic data. The library includes a transformer text processor for generating contextual vector embeddings and a dedicated inference engine for managing large language models. The project distinguishes itself through its ability to process heterogeneous data types, including text, audio, and images, within a unified vision-langu

    Scala
    عرض على GitHub↗4,135
  • hanxiao/bert-as-serviceالصورة الرمزية لـ hanxiao

    hanxiao/bert-as-service

    12,831عرض على GitHub↗

    This project is a high-performance BERT embedding service and inference server designed to map text sequences into fixed-length numerical vectors. It functions as a machine learning microservice and distributed model server that decouples request handling from heavy computation. The system utilizes a ZeroMQ messaging infrastructure to provide low-latency communication between distributed clients and the inference server. It incorporates server-side batch processing and GPU workload scaling to maximize hardware utilization and manage high request volumes. The platform supports semantic search

    Python
    عرض على GitHub↗12,831
  • عرض جميع البدائل الـ 30 لـ Text Embeddings Inference→