24 रिपॉजिटरी
Converting 2D pose detections or images into three-dimensional spatial coordinates.
Distinct from 3D Pose Estimation: Existing candidates are either in 'awesome-lists' or too specific to real-time tracking; a core ML category is needed.
Explore 24 awesome GitHub repositories matching artificial intelligence & ml · 3D Pose Estimation. Refine with filters or upvote what's useful.
SadTalker is a generative framework designed to synthesize expressive talking head videos from static portrait images. By mapping audio signals or text prompts to three-dimensional facial motion coefficients, the system synchronizes lip movements, facial expressions, and head orientation to create realistic digital character performances. The project distinguishes itself by decoupling identity from dynamic motion through latent space encoding, ensuring that the generated animations maintain visual fidelity to the source portrait. It supports comprehensive motion synthesis, including full-body
Calculates head orientation and movement parameters from source data to drive realistic spatial transformations of the static portrait.
AlphaPose एक डीप लर्निंग पोज़ एस्टिमेशन फ्रेमवर्क और PyTorch कंप्यूटर विज़न लाइब्रेरी है, जिसे इमेज और वीडियो में मानव शरीर, चेहरे, हाथ और पैर के कीपॉइंट्स को डिटेक्ट और ट्रैक करने के लिए डिज़ाइन किया गया है। यह स्केलेटल पोस्चर एस्टिमेशन और मल्टी-पर्सन पोज़ ट्रैकिंग के लिए एक सिस्टम प्रदान करता है। यह प्रोजेक्ट थ्री-डायमेंशनल ह्यूमन पोज़ रिकंस्ट्रक्शन के लिए टूल्स लागू करता है, जो टू-डायमेंशनल इमेज डेटा से जॉइंट पोजीशन्स और बॉडी मेश शेप्स जनरेट करता है। इसमें एक मल्टी-पर्सन पोज़ ट्रैकर भी शामिल है जो लगातार वीडियो फ्रेम्स में कई लोगों की पहचान बनाए रखने में सक्षम है। यह फ्रेमवर्क कंप्यूटर विज़न क्षमताओं की एक विस्तृत श्रृंखला को कवर करता है, जिसमें मल्टी-पर्सन कीपॉइंट लोकलाइज़ेशन, ह्यूमन मोशन ट्रैकिंग और थ्री-डायमेंशनल बॉडी मेश का रिकंस्ट्रक्शन शामिल है।
Calculates three-dimensional joint positions and body mesh shapes from two-dimensional image data.
OpenFace is an affective computing framework and facial behavior analysis toolkit designed for the extraction of facial features and the recognition of muscle movements to analyze human emotional behavior. It provides a research platform for analyzing human emotion and social interaction patterns through computer vision. The software implements a suite of tools for detecting facial landmarks, calculating head pose in 3D space relative to a camera, and tracking eye gaze by analyzing eye position and orientation. It also includes capabilities for facial action unit recognition to identify speci
Calculates the orientation and position of a head in 3D space relative to the camera.
MMPose is a PyTorch-based pose estimation toolbox and deep learning training pipeline designed for detecting 2D and 3D keypoints on humans, animals, and faces. It serves as a computer vision model zoo and a framework for both 2D pose estimation and 3D pose lifting. The project is distinguished by its modular architecture and extensibility, employing a registry-based system and hierarchical configurations to allow for custom algorithm integration and model pipeline customization. It supports diverse estimation paradigms, including top-down, bottom-up, and two-stage pose lifting workflows. The
Converts two-dimensional pose detections into three-dimensional coordinates to provide spatial depth information.
DensePose is a 3D human pose estimation framework designed to map 2D image pixels to a 3D surface-based model of the human body in real time. It functions as a computer vision anatomical mapper that projects 2D visual data onto a 3D surface to create detailed anatomical representations. The system operates as an image-to-3D texture transfer engine, localizing 2D image annotations onto 3D models to apply photographic textures to digital human representations. It uses a surface-based body mapping method to associate human pixels in an RGB image with specific coordinates on a 3D body template.
Analyzes 2D images to determine the precise 3D position and orientation of human body parts.
Relativ is an open-source project for the development of custom virtual reality hardware, encompassing the mechanical design, electronics, and software interfaces required to build a headset from scratch. It provides the frameworks necessary for assembling devices using open-source electronics and firmware. The project integrates custom hardware with SteamVR through driver-based configurations, mapping device identifiers and display viewports to ensure rendered images align with physical secondary displays. It employs a combination of microcontroller-based inertial measurement unit polling fo
Uses camera feeds and neural networks to estimate 3D body position for spatial movement tracking.
OpenCVSharp is a .NET library that wraps native OpenCV functions, providing C# developers with access to OpenCV's computer vision capabilities through an API that mirrors the native C/C++ style. It serves as a managed wrapper for image processing, feature detection, object detection, and image manipulation tasks, while also handling automatic disposal of unmanaged OpenCV resources like Mat objects to prevent memory leaks in .NET applications. The library enables keypoint detection and descriptor extraction using algorithms such as AKAZE, BRISK, or FAST, with brute-force or FLANN-based matchin
Estimates the pose of a calibrated camera from a set of 3D-to-2D point correspondences.
SAM 3D Objects is a promptable foundation model that recovers 3D objects and human meshes from single images. It converts masked objects in a single photograph into full 3D models with pose, shape, texture, and layout, while also producing complete 3D human body meshes from the same input. The system integrates promptable segmentation to isolate objects and humans before reconstruction, then aligns the independently reconstructed 3D elements into a shared coordinate space. This enables scene-level understanding where multiple 3D reconstructions from the same image coexist in a common coordina
Estimates 3D shape, pose, and texture from a single image using learned parametric models.
DeepLabCut मार्करलेस 2D और 3D एनिमल पोज़ एस्टिमेशन के लिए एक डीप लर्निंग टूलकिट है। यह एक मोशन ट्रैकिंग सिस्टम के रूप में कार्य करता है जो भौतिक मार्कर्स की आवश्यकता के बिना वीडियो अनुक्रमों में जानवरों पर शारीरिक की-पॉइंट्स की पहचान करता है। यह फ्रेमवर्क विभिन्न प्रजातियों के लिए नेटवर्क्स के प्रशिक्षण में तेज़ी लाने के लिए ट्रांसफर लर्निंग और प्री-ट्रेंड वेट्स की एक लाइब्रेरी का उपयोग करता है। यह वीडियो अनुक्रमों में अद्वितीय पहचान बनाए रखने के लिए मल्टी-इंडिविजुअल आइडेंटिटी ट्रैकिंग को सपोर्ट करता है और लाइव वीडियो फ़ीड्स के लिए रियल-टाइम पोज़ डिटेक्शन प्रदान करता है। यह सिस्टम 3D स्थानिक गति विश्लेषण, शारीरिक मार्कर ट्रैकिंग और सेल्फ-सुपरवाइज्ड प्रेडिक्शन रिफाइनमेंट सहित कंप्यूटर विज़न क्षमताओं की एक विस्तृत श्रृंखला को कवर करता है। इसमें प्रशिक्षण डेटा लेबलिंग के लिए यूटिलिटीज और कस्टम न्यूरल नेटवर्क आर्किटेक्चर को इंटीग्रेट करने के लिए एक मॉडल रजिस्ट्री शामिल है। यह सॉफ़्टवेयर विभिन्न ऑपरेटिंग सिस्टम्स में सुसंगत इंस्टॉलेशन और निष्पादन सुनिश्चित करने के लिए Docker कंटेनराइज़ेशन प्रदान करता है।
Extracts three-dimensional spatial coordinates from video to analyze animal movement in 3D space.
Kalidokit is a web-based motion capture tool that transforms real-time webcam video into 3D character animation data. It functions as a blendshape and kinematics calculator, converting facial, hand, and body tracking data from Mediapipe and TensorFlow.js into blendshape weights and euler rotations for driving digital puppets and avatars. The tool solves face landmarks to derive head rotation, eye blinks, mouth shapes, and brow values for rigging, while hand landmarks are converted into finger joint rotations and body keypoints into per-joint euler rotations for full-body animation. It include
Derives head rotation, tilt, and lean angles from face landmarks using perspective-n-point algorithms.
PRNet is a Python library for 3D facial reconstruction. It uses a deep learning regression model to predict 3D facial geometry and vertex colors from a single 2D input image to generate a textured mesh. The project provides tools for digital face swapping, allowing the replacement of a target face with a new image and blending textures to match the original pose. It also includes a framework for face texture swapping and blending to fit specific 3D poses. Additional capabilities cover facial analysis, including the detection and alignment of facial landmarks and the estimation of head pose a
Implements techniques for calculating the 3D orientation and position of a head relative to a camera.
AniPortrait एक AI वीडियो सिंथेसिस पाइपलाइन है जिसे फ़ोटोरियलिस्टिक स्पीकिंग पोर्ट्रेट और चेहरे के एनिमेशन जनरेट करने के लिए डिज़ाइन किया गया है। यह एक टॉकिंग हेड जनरेटर और ऑडियो-संचालित एनिमेटर के रूप में कार्य करता है जो होंठों की गतिविधियों, अभिव्यक्तियों और सिर के पोज़ को भाषण या संदर्भ वीडियो स्रोतों के साथ सिंक्रोनाइज़ करता है। सिस्टम में एक स्थिर संदर्भ छवि पर स्रोत वीडियो से गतिविधियों को फिर से लागू करने के लिए एक चेहरे की अभिव्यक्ति स्थानांतरण टूल शामिल है। यह जनरेट किए गए फ़्रेमों में दृश्य पहचान और स्थिरता बनाए रखने के लिए संदर्भ-आधारित छवि कंडीशनिंग के साथ एक लेटेंट डिफ्यूजन मॉडल का उपयोग करता है। पाइपलाइन ऑडियो-टू-एक्सप्रेशन मैपिंग, पोज़-गाइडेड मोशन कंट्रोल और फ़ोटोरियलिस्टिक वीडियो सिंथेसिस को कवर करती है। यह जनरेशन प्रक्रिया में तेज़ी लाने और कुल रेंडरिंग समय को कम करने के लिए फ़्रेम इंटरपोलेशन अपसैंपलिंग को शामिल करती है।
Directs specific head orientation and movement during animation using external control files.
This project is a scientific computing framework for the .NET ecosystem, providing a comprehensive suite of libraries for numerical analysis, statistics, and mathematical optimization. It serves as a foundational toolkit for developing applications in machine learning, digital signal processing, and computer vision. The framework provides specialized toolkits for training and deploying predictive models, including neural networks, support vector machines, and decision trees. It further distinguishes itself with deep integrations for real-time visual analysis, such as object tracking and facia
Calculates the three-dimensional position and orientation of objects in space.
GET3D is a generative 3D mesh model and rendering framework designed to synthesize high-quality textured shapes and tetrahedral meshes. It functions as an image-to-3D reconstructor and text-to-3D generator, utilizing a differentiable 3D renderer to produce realistic visual perspectives and material effects. The system enables the creation of 3D assets from single 2D images, point clouds, or descriptive text prompts. It features a latent space interpolator for creating smooth transitions between different 3D objects and supports the independent control of geometry and texture. The project cov
Estimates 3D geometry, textures, and lighting from a single-view image without requiring 3D supervision data.
EasyMocap is a markerless 3D human motion capture system that recovers body, hand, and face poses from single or multi-view video without physical markers or suits. It uses parametric body models like SMPL, SMPL-X, and MANO, and leverages mirror reflections to resolve depth ambiguity in single-view pose estimation, improving accuracy by computing mirror surface normals from vanishing points. The system distinguishes itself through mirror-assisted depth disambiguation, enabling accurate 3D pose reconstruction from a single RGB image or video that includes a mirror reflection. It also supports
Recovering 3D body, hand, and face poses from single or multi-view video without physical markers or suits.
lite.ai.toolkit एज AI तैनाती के लिए डिज़ाइन किया गया एक C++ कंप्यूटर विज़न टूलकिट है। यह संसाधन-सीमित उपकरणों पर ऑब्जेक्ट डिटेक्शन, इमेज क्लासिफिकेशन और सेगमेंटेशन के लिए प्री-ट्रेंड मॉडल के निष्पादन को सक्षम बनाता है। इस प्रोजेक्ट में एक मल्टी-बैकएंड इन्फरेंस इंजन है जो ONNX मॉडल रनटाइम का समर्थन करता है, जिससे AI मॉडल को विभिन्न हार्डवेयर लक्ष्यों पर चलने की अनुमति मिलती है। इसमें लेटेंसी को कम करने और प्रोसेसिंग गति बढ़ाने के लिए विशेष रूप से NVIDIA हार्डवेयर के लिए एक GPU-त्वरित पाइपलाइन शामिल है। यह टूलकिट चेहरे के विश्लेषण की क्षमताओं की एक विस्तृत श्रृंखला को कवर करता है, जिसमें भावना पहचान, लिंग और आयु अनुमान और हेड पोज़ विश्लेषण शामिल है। यह फीचर एम्बेडिंग के निष्कर्षण और पहचान को सत्यापित करने के लिए कोसाइन समानता की गणना के माध्यम से चेहरे की पहचान के लिए उपकरण भी प्रदान करता। अतिरिक्त क्षमताओं में फोरग्राउंड आइसोलेशन के लिए इमेज मैटिंग, ग्रेस्केल इमेज कलराइज़ेशन और आर्टिस्टिक स्टाइल ट्रांसफर शामिल हैं।
Calculates 3D head orientation using yaw, pitch, and roll Euler angles.
Depth-Anything-3 is a collection of core model implementations for depth prediction, multi-view geometry estimation, and RGB-D spatial pipelines. It includes a monocular depth estimation model for predicting depth maps from single images or video, and a 3D Gaussian splatting generator that predicts parameters to synthesize high-fidelity novel views of a scene. The project provides a multi-view geometry estimator for calculating spatially consistent depth and camera poses across synchronized visual inputs. It also functions as a visual SLAM enhancement tool designed to reduce drift and improve
Calculates precise 3D camera positions and orientations from visual inputs using coordinate-based pose estimation.
evo SLAM एल्गोरिदम, रोबोट ओडोमेट्री और प्रक्षेपवक्र डेटा के मूल्यांकन के लिए एक Python फ्रेमवर्क है। यह अनुमानित पथों और ग्राउंड ट्रुथ संदर्भों के बीच पूर्ण और सापेक्ष पोज़ त्रुटियों की गणना करके ड्रिफ्ट और सटीकता को मापने के लिए एक विश्लेषण लाइब्रेरी के रूप में कार्य करता है। यह प्रोजेक्ट स्थानिक प्रक्षेपवक्रों के बीच रोटेशन, ट्रांसलेशन और स्केल को सही करने के लिए एक ज्यामितीय संरेखण ढांचा प्रदान करता है, जो सुसंगत त्रुटि माप सुनिश्चित करता है। इसमें ओडोमेट्री ड्रिफ्ट विश्लेषण और रोबोटिक्स डेटा के प्रसंस्करण के लिए विशेष उपकरण शामिल हैं, जिसमें ROS बैगफाइल्स से प्रक्षेपवक्र जानकारी निकालने की क्षमता भी शामिल है। यह सॉफ़्टवेयर भौगोलिक मानचित्र टाइल्स और ROS मानचित्र ओवरले के समर्थन के साथ 2D और 3D प्रक्षेपवक्र विज़ुअलाइज़ेशन सहित क्षमताओं की एक विस्तृत श्रृंखला को कवर करता है। अतिरिक्त कार्यक्षमता में टाइमस्टैम्प सिंक्रोनाइज़ेशन, स्थानिक परिवर्तन और विभिन्न उद्योग-मानक प्रारूपों में प्रक्षेपवक्र डेटा को फ़िल्टर या निर्यात करने की क्षमता शामिल है।
Flattens 3D trajectory poses into specified 2D planes such as xy, xz, or yz for simplified analysis.
VideoPose3D is a machine learning framework designed for 3D human pose estimation. It functions as a motion reconstruction tool that predicts 3D joint positions from 2D video sequences using a temporal convolutional network to process body movement over time. The project includes a semi-supervised learning pipeline that improves pose accuracy by combining labeled datasets with unlabeled video data and projection consistency loss. It also features a video pose visualizer capable of rendering 3D skeleton reconstructions and 2D keypoints as overlays on original footage. The framework covers the
Provides a framework for converting 2D pose detections from video into three-dimensional spatial coordinates.
This project is a 3D visual localization framework designed to determine a camera's exact position and orientation by matching 2D image features against a 3D reference model. It includes a structure-from-motion pipeline to reconstruct 3D scene geometry from unordered image sets, creating the necessary spatial maps for localization. The system employs a hierarchical coarse-to-fine localization approach. This process begins with a global-descriptor image retrieval system to identify candidate reference images from a large database and progresses through local feature matching to final 3D model
Calculates precise camera position and orientation by solving the transformation between 2D points and 3D coordinates.