21 रिपॉजिटरी
Tools for managing and utilizing extended context windows during inference requests.
Distinct from Long Context Training Optimizations: Distinct from Long Context Training Optimizations: focuses on inference-time context utilization, not training-time memory management.
Explore 21 awesome GitHub repositories matching artificial intelligence & ml · Context Window Management. Refine with filters or upvote what's useful.
Embedchain is an LLM memory management framework and RAG orchestration engine designed to provide AI agents with a persistent storage layer. It functions as a long-term memory pipeline that extracts facts from unstructured interactions and stores them as permanent knowledge base entries to retain user preferences and interaction history across sessions. The system employs a hybrid vector database interface that combines semantic embeddings with traditional keyword search. It utilizes an entity-linking knowledge graph to connect related information points and applies temporal ranking to distin
Partitions context into distinct user, session, and agent layers to maintain state across interaction scales.
Sglang is a high-performance inference engine and serving system designed for large language and multimodal models. It provides a programmable interface for orchestrating complex generation workflows, enabling developers to coordinate multi-turn dialogues, tool invocations, and reasoning chains through a domain-specific language. The platform is built to support production-scale deployments, offering an OpenAI-compatible API that allows for integration with existing application ecosystems. The system distinguishes itself through a disaggregated architecture that separates compute-intensive pr
Utilizes extended context windows to ingest and reason over large documents during inference requests.
This project provides a framework for managing multi-agent systems, designed to automate complex software development, infrastructure, and business workflows. It functions as a multi-agent workflow orchestrator that routes tasks to domain-specific workers while maintaining state persistence and infrastructure automation. By leveraging large language models, the system decomposes high-level objectives into actionable plans, ensuring that complex operations are executed with consistency and reliability. The framework distinguishes itself through its hierarchical agent registry and policy-driven
Manages memory and information prioritization to maximize efficiency within limited context windows.
Qwen-7B is a pretrained causal language model designed for natural language generation, text processing, and complex reasoning tasks. It is available as an instruction-tuned model optimized for conversational interactions and a tool-use model capable of executing function calls and interacting with external APIs. The project provides a quantized version of the model to reduce GPU memory usage and supports the development of autonomous agents that can execute code and perform functions to complete complex goals. The system covers a wide range of capabilities including model fine-tuning throug
Extends the model's context window by scaling position indices of the input sequence.
Qwen is a comprehensive framework for large language model development, serving, and deployment. It provides a complete ecosystem for transformer-based sequence modeling, offering base models alongside specialized tools for instruction-tuned alignment, fine-tuning, and long-context inference. The project is designed to support both research and production environments, enabling users to train, optimize, and host generative models locally or across distributed hardware. The framework distinguishes itself through its focus on high-performance serving and extensibility. It features a high-perfor
Optimizes serving environments to handle extended token windows and long-context sequences using advanced attention scaling.
Qwen2-VL is a multimodal large language model and vision language model designed to process and reason across text, images, and video content. It functions as a visual reasoning engine and a visual agent framework, capable of interpreting visual data to perform object detection, document parsing, and spatial reasoning. The model is distinguished by its ability to act as a video understanding model, processing hour-long videos with second-level indexing and event recall. It further differentiates itself through a visual agent capability that interacts with software interfaces and robotic hardw
Increases capacity for ultra-long documents and videos by scaling position embeddings to extend the context window.
Code Llama is a large language model based on Llama 2 trained specifically for programming tasks and software development. It provides specialized model types optimized for general code generation, instruction following, and context-aware infilling. The project includes an instruction-tuned programming model for executing technical tasks via natural language prompts and a code infilling model that predicts missing sections based on surrounding source context. A large context code model is also provided to analyze extensive blocks of source code for improved coherence. The system covers capab
Extends the maximum token processing limit by adjusting and scaling positional embeddings.
This project is an AI agent workflow orchestrator and automated software lifecycle manager designed to sequence specialized AI personas for end-to-end software development. It serves as a prompt engineering library and a full-stack development toolkit that guides the process from initial discovery and specification through to deployment and code review. The system features a context management framework that utilizes progressive loading and routing tables to fetch reference files on-demand, reducing token consumption within the model context window. It employs a definition-based routing syste
Optimizes LLM token usage through progressive loading and routing of reference files.
LMFlow is a comprehensive suite for large language model fine-tuning, context extension, multimodal processing, and inference execution. It provides a toolkit for updating model parameters through full tuning or memory-efficient adapter algorithms, alongside an inference engine for executing tuned models via command-line or web-based interfaces. The framework includes a dedicated alignment suite for supervised tuning and reward model training to refine model behavior. It features a context window extender to increase maximum input lengths and a multimodal framework for building chatbots that
Increases maximum input sequence lengths by scaling positional embeddings.
This project provides a Chinese large language model based on the LLaMA architecture. It is an instruction-tuned model optimized for natural language processing and multi-turn conversations in Chinese. The system includes a framework for parameter-efficient fine-tuning using low-rank adaptation and quantization to reduce memory requirements. It also implements retrieval augmented generation for local document question answering and supports long-context processing for sequences up to 64K tokens. The project covers a broad set of capabilities including supervised instruction tuning, reinforce
Extends the context window by scaling position embeddings to handle longer sequences without retraining.
GLM-4 is an open weights large language model designed as a multimodal chat system. It functions as a reasoning-focused and multilingual model capable of processing and generating responses across text and visual data types. The model is distinguished by its function-calling capabilities, allowing it to interface with external tools and APIs to execute tasks and retrieve real-time information. It is optimized for complex logical reasoning, mathematical problem solving, and deep research involving long-form content generation. Broad capabilities include multilingual text generation, the creat
Expands the maximum token limit for long-form inputs and outputs by scaling positional embeddings.
pi-autoresearch is an autonomous research extension that automates iterative code-editing and performance-measurement loops driven by large language models. It functions as an experiment lifecycle automator, executing repetitive cycles of changes and benchmarks until a specific goal is reached. The system distinguishes itself by organizing successful experimental trials into independent git branches for review and merging. It includes a real-time research dashboard for monitoring metrics and status, and utilizes median absolute deviation to calculate confidence scores that filter benchmark no
Manages LLM state through automated summaries and context injection to prevent information loss during long sessions.
GLM-4 is a large language model and fine-tuning framework designed for human-like text production, complex reasoning, and multilingual conversation. It functions as a multimodal system capable of processing high-resolution visual content and as a long-context model designed to analyze documents with a context window of up to one million tokens. The project differentiates itself through a function calling interface that enables AI agent development by connecting the model to external APIs and real-time web browsing. It includes specialized capabilities for generating functional programming cod
Extends the token window using positional embedding scaling to process up to one million tokens.
This is an open-source Python SDK for building and orchestrating production-grade AI agents. It provides a unified framework for creating conversational agents that can use tools, maintain state, and coordinate across multiple language model providers including OpenAI, Anthropic, Google, Amazon Bedrock, and locally-hosted models. The SDK supports multi-agent orchestration through graphs, teams, and swarms, allowing several specialized agents to collaborate on complex tasks. Agents can be composed as callable tools that other agents invoke, and the framework includes policy handlers that inspe
Keeps conversation history within a token budget by automatically discarding older turns.
Shimmy एक स्थानीय लार्ज लैंग्वेज मॉडल इन्फरेंस इंजन और सर्वर है जो GGUF फॉर्मेट वाले वेट्स को लोड और सर्व करता है। इसे Rust में लिखे गए एक सिंगल बाइनरी रनटाइम के रूप में वितरित किया जाता है, जो बाहरी रनटाइम डिपेंडेंसीज के बिना मॉडल चलाने के लिए एक स्टैंडअलोन एनवायरनमेंट प्रदान करता है। यह प्रोजेक्ट हार्डवेयर एक्सेलेरेशन के लिए WebGPU का उपयोग करता है, जिससे मॉडल कंप्यूट कर्नेल एक मानकीकृत इंटरफेस के माध्यम से विविध ग्राफिक्स हार्डवेयर पर निष्पादित हो सकते हैं। इसमें एक स्थानीय सर्वर है जो OpenAI-संगत API लेयर को लागू करता है, जिससे एप्लिकेशन्स मानकीकृत REST एंडपॉइंट्स के माध्यम से स्थानीय मॉडल्स के साथ इंटरफेस कर सकते हैं। मेमोरी और प्रदर्शन को GPU VRAM उपयोग को कम करने के लिए क्वांटाइज्ड की-वैल्यू कैश कंप्रेशन और मॉडल कॉन्टेक्स्ट विंडो को बढ़ाने के लिए रोटरी एम्बेडिंग स्केलिंग के माध्यम से मैनेज किया जाता है। सिस्टम में स्थानीय स्टोरेज से संगत वेट्स को स्कैन और रजिस्टर करने के लिए ऑटोमैटिक मॉडल फाइल डिस्कवरी भी शामिल है। सर्वर को ऑपरेशंस को नियंत्रित करने और मॉडल जनरेशन को सत्यापित करने के लिए एक समर्पित कमांड-लाइन इंटरफेस के माध्यम से मैनेज किया जाता है।
Extends the model context window by scaling the frequency of positional embeddings.
bert4keras Keras डीप लर्निंग फ्रेमवर्क के लिए BERT ट्रांसफॉर्मर आर्किटेक्चर का एक हल्का पुनर्कियान्वयन है। यह टेक्स्ट क्लासिफिकेशन, सीक्वेंस लेबलिंग, और सिमेंटिक एम्बेडिंग एक्सट्रैक्शन के लिए उपयोग किया जाने वाला एक नेचुरल लैंग्वेज प्रोसेसिंग टूलकिट और ट्रांसफॉर्मर मॉडल लाइब्रेरी के रूप में कार्य करता है। इस फ्रेमवर्क में प्रश्न उत्तर और टेक्स्ट जनरेशन के लिए एक सीक्वेंस-टू-सीक्वेंस मॉडल सिस्टम शामिल है, साथ ही रियल-टाइम भविष्यवाणियों के लिए प्रशिक्षित ट्रांसफॉर्मर्स को वेब APIs के रूप में तैनात करने के लिए एक मॉडल इन्फरेंस सर्वर भी शामिल है। क्षमताएं नेचुरल लैंग्वेज अंडरस्टैंडिंग कार्यों की एक विस्तृत श्रृंखला को कवर करती हैं, जिसमें रीडिंग कॉम्प्रिहेंशन, रिलेशन एक्सट्रैक्शन और लॉन्ग टेक्स्ट प्रोसेसिंग शामिल है। यह लाइब्रेरी पैरामीटर रिडक्शन, मजबूती के लिए एडवर्सरियल ट्रेनिंग, और लेयर-वाइज लर्निंग रेट कॉन्फ़िगरेशन जैसी ऑप्टिमाइज़ेशन तकनीकों के साथ-साथ भाषा मॉडल प्री-ट्रेनिंग और फाइन-ट्यूनिंग के लिए टूल्स प्रदान करती है। इस प्रोजेक्ट में बाहरी फॉर्मेट्स से प्रशिक्षित वेट्स को संगत Keras स्ट्रक्चर्स में बदलने के लिए एक वेट-कन्वर्जन लोडर शामिल है।
Provides hierarchical position embeddings to handle text sequences that exceed standard transformer length limits.
LLaVA-NeXT एक मल्टीमॉडल लार्ज लैंग्वेज मॉडल फ्रेमवर्क और ट्रेनिंग टूलकिट है जिसे टेक्स्ट जनरेट करने के लिए इंटरलीव्ड इमेजेस और वीडियो सीक्वेंस को प्रोसेस करने के लिए डिज़ाइन किया गया है। यह एक विजुअल लैंग्वेज मॉडल के रूप में कार्य करता है जो जटिल तर्क, प्रश्न उत्तर, और वीडियो समझ को निष्पादित करने के लिए विजन एनकोडर्स को लैंग्वेज मॉडल्स के साथ जोड़ता है। यह सिस्टम घटनाओं का वर्णन करने, कार्यों का सारांश देने, और कई विजुअल इनपुट्स में तर्क करने के लिए हाई-रिज़ॉल्यूशन इमेजेस और टेम्पोरल वीडियो फ्रेम्स का विश्लेषण करने में सक्षम है। यह डॉक्यूमेंट्स और चार्ट्स की व्याख्या, स्थानिक वातावरण विश्लेषण, और इमेजेस और वीडियो दोनों के लिए वर्णनात्मक कैप्शन के जनरेशन को सपोर्ट करता है। इस फ्रेमवर्क में मतिभ्रम (hallucinations) को कम करने और सटीकता में सुधार करने के लिए प्रिफरेंस ऑप्टिमाइज़ेशन के माध्यम से मल्टीमॉडल मॉडल्स को ट्यून करने के लिए टूल्स शामिल हैं। यह इन क्षमताओं को HTTP बैकएंड के माध्यम से API सर्विस के रूप में डिप्लॉय करने के लिए एक इन्फरेंस सर्वर भी प्रदान करता है।
Scales position embeddings via interpolation to extend the context window for longer video sequences.
OpenSquilla is an LLM agent orchestration framework designed to coordinate multi-step AI workflows and tool execution using directed acyclic graphs. It functions as a centralized system for managing specialized skill packages and executing complex reasoning sequences. The project distinguishes itself through a routing gateway that directs tasks to different AI providers based on complexity, cost, and performance. It utilizes a multi-tier AI memory system that organizes working, episodic, and semantic knowledge using local embeddings and SQLite, alongside a secure execution sandbox that isolat
Deno AI Agent triggers manual compaction of long sessions to maintain token efficiency and context window limits.
Positron is a data science integrated development environment and AI-powered code editor designed for polyglot development, specifically supporting Python and R. It functions as a remote compute workspace that separates the user interface from the execution kernel via SSH or container integration. The environment features a deep integration of large language models that provide context-aware suggestions and automated data analysis by accessing real-time interpreter state, in-memory objects, and plot outputs. It distinguishes itself through a polyglot runtime bridge that enables cross-language
Implements token compaction and metadata management to keep long conversations within the model's context window.
Langroid is a multi-agent orchestration framework and tool integration suite designed for building complex AI applications. It serves as a multi-modal integration layer that connects diverse local and remote language models with an agentic retrieval-augmented generation system. The project distinguishes itself through a collaborative message-exchange paradigm, allowing specialized agents to delegate tasks hierarchically and coordinate via structured communication. It features an advanced state management system for conversational AI, including the ability to rewind and prune conversation hist
Manages and expands the context window by coalescing overlapping document chunks around search matches.