12 个仓库
Software designed to automatically identify and categorize objects within digital images.
Distinguishing note: The candidates were either raw image developers or generative AI, not classification systems.
Explore 12 awesome GitHub repositories matching artificial intelligence & ml · Image Recognition Systems. Refine with filters or upvote what's useful.
Darknet is a low-level neural network engine and framework written in C. It is designed for training and deploying deep learning models, with a primary focus on convolutional neural networks. The project serves as a CUDA accelerated deep learning library that offloads heavy mathematical operations to NVIDIA graphics hardware. This acceleration is used to increase processing speed and reduce execution time during the training of large networks. The engine supports a range of activities including deep learning research, image recognition development, and the training of convolutional neural ne
Enables the development of systems that automatically identify and categorize objects within images.
This project is an automated machine learning framework and toolkit designed for training and tuning custom models for classification, regression, and recommendations. It functions as a multimodal machine learning toolkit capable of processing and training models using a combination of text, image, audio, and sensor data. The framework distinguishes itself as a multimodal data processor that can handle and visualize large datasets on a single machine using column-oriented disk storage. It includes a core machine learning model generator that converts trained models into formats compatible wit
Trains models to classify visual content, detect objects with bounding boxes, and identify visually similar images.
ConvNetJS is a JavaScript deep learning library and neural network training engine designed for client-side machine learning. It functions as a framework for building, training, and running convolutional neural networks directly within a web browser without the need for a backend server. The library specializes in image recognition and pattern analysis using convolutional and pooling layers. It enables the creation of models for classification and regression tasks, as well as the development of reinforcement learning agents that optimize behavior through trial and error in simulated environme
Identifies and categorizes objects and visual features within digital images using convolutional neural networks.
KoboldCPP is a local large language model inference engine and GGUF model runner designed to execute quantized models on personal hardware. It functions as a multimodal AI server and API gateway, providing OpenAI-compatible endpoints that allow third-party clients to interact with locally hosted models. The project distinguishes itself as an AI storytelling backend, featuring dedicated tools for long-form narrative management through persistent memory, world lore tracking, and character state management. It further extends its capabilities as a multimodal server capable of processing text, im
Analyzes visual inputs to describe or interpret images using multimodal vision capabilities.
This project is a Python implementation of the Faster R-CNN object detection framework. It serves as a convolutional neural network library and tool for locating and classifying multiple objects within images. The framework provides a pre-trained model implementation that allows for object detection inference without manual training. It supports the full lifecycle of object detection, including training detectors on visual datasets to identify and bound specific object classes. The system covers capabilities for computer vision model evaluation, neural network optimization to reduce model si
Implements a system to automatically identify and categorize multiple objects within digital images.
Anti-Anti-Spider is an automated web scraping toolkit and CAPTCHA bypass framework. It uses convolutional neural networks to recognize characters and digits in image-based security challenges, enabling programmatic access to protected web content. The project functions as an image recognition model trainer, providing a workflow to preprocess labeled image datasets and train custom neural networks. Users can configure model architectures and hyperparameters to align the recognition system with the visual style of specific target websites. The toolkit covers capabilities for image data preproc
Trains neural networks to automatically identify and categorize characters within custom image datasets.
The TensorFlow Cookbook is a collection of code examples and recipes for building, training, and deploying machine learning models using TensorFlow. It covers the full model lifecycle, from constructing neural networks and training them with configurable parameters to packaging trained models for production deployment with unit tests and multi-device support. The project also integrates TensorBoard for logging and visualizing computational graphs, scalar summaries, and histograms during training. The cookbook demonstrates a wide range of machine learning techniques, including convolutional ne
Applies convolutional neural networks to classify images, retrain architectures, and generate artistic effects.
is-thirteen 是一个数字验证库和数值相等性检查器,旨在验证给定输入是否等于十三。它充当数据分类工具,可识别跨数值、文本和视觉输入流的这一特定值。 该项目包含一个基于图像的数字分类器,使用深度学习和神经网络分析来识别上传图像中数字十三的视觉表现。 该库涵盖了多种验证方法,包括精确算术相等性、定义容差范围内的近似值匹配、科学计数法解析以及书写形式的语言模式匹配。
Automatically identifies and categorizes the number thirteen within digital images.
This is an open-source automation tool for the game Wuthering Waves that uses image recognition to control gameplay without modifying game memory or files. It runs automation tasks while the game window is minimized or obscured, freeing the computer for other use, and accepts command-line arguments to start specific tasks and optionally exit after completion. The tool automatically detects playable characters through screen analysis and adapts actions without manual skill configuration. It supports all common 16:9 resolutions up to 4K as well as some ultrawide formats, with a minimum required
Maa simulates user inputs by analyzing screen images to automate game interactions without memory or file modification.
PyBoy 是一个用 Python 编写的可编程 Game Boy 模拟器和硬件仿真框架。它作为一个仿真引擎,允许用户执行原始掌机软件,同时提供程序化接口来控制、探测和自动化游戏执行。 该项目专门设计为强化学习环境,暴露模拟器状态和控制以促进机器学习智能体的训练。其特色在于提供游戏区域映射工具,以及提取简化的 2D 屏幕表示和碰撞图,以支持人工智能。 该系统涵盖了广泛的功能,包括周期精确的硬件仿真、直接内存读写操作以及用于执行钩子的回调系统。它支持提取实时游戏数据(如精灵位置和内存符号),并包含一个通过绕过图形和音频渲染来加速模拟速度的无头执行模式。 该模拟器还提供通过快照序列化进行状态持久化的工具、用于自主智能体的输入模拟,以及用于内存分析和 ROM 数据修改的工具。
Enables automated gameplay and behavior verification through scripted inputs and memory state monitoring.
该项目是 Faster R-CNN 目标检测架构的 PyTorch 实现。它提供了一个框架,用于使用深度学习系统识别图像中的多个对象类别及其对应的边界框。 该实现包括用于在自定义数据集上优化模型的训练流水线,以及用于将预训练权重从外部格式转换为模型初始化兼容结构的工具。 该系统涵盖了包含区域建议网络(RPN)和 ROI 池化层的两阶段检测流水线。它结合了多任务损失函数和基于锚点的边界框回归来细化对象位置。 该项目包含用于实时可视化训练损失和预测准确率的工具,以监控模型性能。
Identifies and categorizes specific items within digital images using trained neural network models.
Recognize-anything is a multimodal foundation model designed for image recognition, visual tagging, and the generation of descriptive text captions from visual input. It functions as a multimodal embedding model that maps images and text into a shared vector space to enable cross-modal retrieval and recognition. The system implements zero-shot image classification and open-vocabulary object detection, allowing it to recognize object categories not present in the original training data through custom label embeddings. It also features a visual tagging engine and a captioning system that produc
Provides a comprehensive system to automatically identify and categorize objects within digital images.