awesome-repositories.com
ब्लॉग
awesome-repositories.com

AI-संचालित खोज के साथ बेहतरीन ओपन-सोर्स रिपॉजिटरी खोजें।

एक्सप्लोर करेंक्यूरेटेड खोजेंओपन-सोर्स विकल्पसेल्फ-होस्टेड सॉफ्टवेयरब्लॉगसाइटमैप
प्रोजेक्टहमारे बारे मेंहम रैंकिंग कैसे करते हैंप्रेसMCP सर्वर
कानूनीगोपनीयताशर्तें
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
R

rchipka/node-osmosis

0
View on GitHub↗
4,110 स्टार्स·260 फोर्क्स·JavaScript·6 व्यूज़

Node Osmosis

यह प्रोजेक्ट एक Node.js वेब स्क्रैपिंग फ्रेमवर्क है जिसे रिक्वेस्ट, पार्सिंग और डॉक्यूमेंट इंटरैक्शन के प्रोग्रामेटिक वर्कफ़्लो के माध्यम से डेटा निष्कर्षण को स्वचालित करने के लिए डिज़ाइन किया गया है। यह एक हेडलेस वेब क्रॉलर, एक HTTP रिक्वेस्ट मैनेजर, और एक DOM पार्सर और एक्सट्रैक्टर के रूप में कार्य करता है।

फ्रेमवर्क डायनामिक सामग्री के साथ बातचीत करने के लिए एक JavaScript निष्पादन इंजन और CSS और XPath सिलेक्टर्स दोनों का उपयोग करने वाले एक हाइब्रिड चयन प्रणाली को जोड़कर खुद को अलग करता है। इसमें प्रमाणित स्थितियों को बनाए रखने और स्वचालित ट्रैफिक को प्रबंधित करने के लिए प्रॉक्सी रोटेशन और कुकी-जार सत्र प्रबंधन के लिए विशेष मिडलवेयर शामिल है।

इसकी व्यापक क्षमताओं में रिकर्सिव लिंक क्रॉलिंग, पेजिनेशन हैंडलिंग और वेब फॉर्म ऑटोमेशन शामिल हैं। टूल ट्रैफिक प्रबंधन सुविधाएं भी प्रदान करता है जैसे कि समयबद्ध देरी के माध्यम से रिक्वेस्ट रेट लिमिटिंग और कस्टम HTTP हेडर कॉन्फ़िगरेशन।

Features

  • Web Crawlers - Functions as a programmable web crawler for asynchronously visiting URLs and extracting structured data.
  • Web Scraping and Extraction - Implements a framework for parsing HTML and using selectors to extract structured data from websites.
  • Web Scraping Frameworks - Offers a comprehensive framework for automating data extraction from websites via programmatic workflows.
  • DOM Tree Construction - Converts raw HTML and XML strings into a hierarchical document object model for programmatic extraction.
  • CSS and XPath Query Engines - Provides a query engine that locates page elements using both CSS3 selectors and XPath 1.0 expressions.
  • HTML Selector Extractors - Implements CSS and XPath selectors to extract structured data from HTML and XML documents.
  • CSS Selector - Extracts specific fields from HTML pages by evaluating CSS selectors against the DOM.
  • HTTP Request Management - Provides comprehensive tools for executing and managing the lifecycle of HTTP requests including custom headers and proxies.
  • XML and HTML Document Parsers - Parses HTML and XML content from strings into searchable in-memory document trees.
  • Cookie-Based Session Management - Manages authenticated sessions by storing and matching cookies with domain and path specificity.
  • Hybrid CSS-XPath Selectors - Retrieves specific elements and attributes using a hybrid of CSS and XPath syntax.
  • HTTP Request Managers - Manages HTTP communication parameters, including cookies, custom headers, and proxy rotation.
  • Web Crawling - Systematically discovers and navigates web content across domains for large-scale data collection.
  • Domain-Restricted Link Navigation - Navigates through linked pages with configurable restrictions on internal and external domains.
  • JavaScript Content Retrievals - Retrieves content from pages where data is generated dynamically via JavaScript by simulating events.
  • Web Page Pagination Discovery - Automatically identifies and follows pagination links to traverse multi-page result sets.
  • Recursive Web Discovery - Automatically discovers new pages by following hyperlinks recursively through website hierarchies.
  • Custom Request Headers - Allows the definition of custom key-value pairs in outgoing HTTP request headers.
  • Proxy Request Routers - Distributes outgoing requests across a pool of proxies to mask the origin IP and avoid bans.
  • Proxy and User-Agent Rotation Middleware - Ships middleware that rotates proxies and user-agent strings to avoid detection during request processing.
  • Proxy Routing - Distributes outgoing requests across multiple proxy servers to avoid IP-based rate limits.
  • DOM Event Execution Environments - Provides an environment to run scripts that interact with dynamic content and trigger DOM events.
  • Element-Targeted JavaScript Execution - Runs JavaScript within the document to interact with dynamic content and trigger DOM events.
  • Randomized Request Delays - Inserts configurable pauses between requests to mimic human browsing patterns and avoid server rate limits.
  • Request Rate Limiting - Introduces timed delays between requests to prevent server overload and evade detection.
  • Dynamic Element Event Handling - Handles event listeners and scripts to interact with elements added to the DOM asynchronously.
  • Programmatic Form Submissions - Triggers web form submissions programmatically by simulating button clicks and handling attributes.
  • Dynamic Content Extraction - Extracts data from JavaScript-heavy pages by simulating browser events and rendering dynamic content.
  • Contextual - Requests content using URLs dynamically constructed from request info, data objects, or context searches.
  • Web Form Filling Tools - Provides tools to programmatically fill and submit web forms to access gated content.
  • JavaScript Crawling Frameworks - HTML/XML parser and scraper for Node.js.
  • Web Scraping - HTML/XML parser and web scraper.

स्टार हिस्ट्री

rchipka/node-osmosis के लिए स्टार हिस्ट्री चार्टrchipka/node-osmosis के लिए स्टार हिस्ट्री चार्ट

AI सर्च

और अधिक बेहतरीन रिपॉजिटरी खोजें

अपनी ज़रूरत को सरल भाषा में बताएं — AI हजारों क्यूरेटेड ओपन-सोर्स प्रोजेक्ट्स को प्रासंगिकता के आधार पर रैंक करता है।

Start searching with AI

Node Osmosis के ओपन-सोर्स विकल्प

समान ओपन-सोर्स प्रोजेक्ट्स, जो Node Osmosis के साथ साझा की गई सुविधाओं के आधार पर रैंक किए गए हैं।
  • bda-research/node-crawlerbda-research का अवतार

    bda-research/node-crawler

    6,785GitHub पर देखें↗

    node-crawler is a programmable web crawler for Node.js that manages request queues and automates data extraction. It functions as a rate-limited HTTP client and a headless HTML parser, providing the infrastructure to visit large sets of URLs asynchronously while preventing duplicate processing through task deduplication. The project distinguishes itself through a proxy rotation manager that cycles user agents and proxy servers to bypass access restrictions. It utilizes the HTTP/2 protocol to improve request performance and server compatibility during large-scale scraping operations. The syst

    TypeScriptcheeriocrawlerextract-data
    GitHub पर देखें↗6,785
  • remitchell/python-scrapingREMitchell का अवतार

    REMitchell/python-scraping

    4,714GitHub पर देखें↗

    This project is a Python web scraping library and automated data collection suite. It provides tools for extracting structured data from websites, implementing web crawlers to navigate site links, and parsing HTML DOM structures to isolate specific elements and attributes. The toolkit includes a pipeline for processing unstructured text and cleaning raw web content to extract meaningful information. It also features capabilities for image data extraction and the integration of external APIs to retrieve structured data from remote endpoints. The system covers broad capability areas including

    Jupyter Notebook
    GitHub पर देखें↗4,714
  • lapwinglabs/x-raylapwinglabs का अवतार

    lapwinglabs/x-ray

    5,904GitHub पर देखें↗

    X-Ray is a web scraping framework and asynchronous web crawler designed to extract structured data from websites. It functions as an HTML data extractor that transforms raw page content into a defined schema using CSS-style selectors. The project implements a headless browser crawler capable of executing JavaScript to render dynamic content. It handles website content discovery through a breadth-first crawling strategy and automatic pagination discovery to traverse multi-page result sets. The framework manages web data pipelines using a concurrency-limited request queue and request rate cont

    JavaScript
    GitHub पर देखें↗5,904
  • symfony/dom-crawlersymfony का अवतार

    symfony/dom-crawler

    4,043GitHub पर देखें↗

    This project is an HTML and XML DOM parser designed for loading and navigating the structure of web documents to extract specific data points. It functions as a web scraping utility that provides a system for locating precise elements using a CSS and XPath selector engine. The library includes a URI resolver that converts relative links found in documents into absolute addresses using a base URI. It provides a set of tools for retrieving text, attributes, and media sources from parsed content. The toolset covers document hierarchy traversal, selector-based filtering, and text extraction with

    PHP
    GitHub पर देखें↗4,043
Node Osmosis के सभी 30 विकल्प देखें→

अक्सर पूछे जाने वाले प्रश्न

rchipka/node-osmosis क्या करता है?

यह प्रोजेक्ट एक Node.js वेब स्क्रैपिंग फ्रेमवर्क है जिसे रिक्वेस्ट, पार्सिंग और डॉक्यूमेंट इंटरैक्शन के प्रोग्रामेटिक वर्कफ़्लो के माध्यम से डेटा निष्कर्षण को स्वचालित करने के लिए डिज़ाइन किया गया है। यह एक हेडलेस वेब क्रॉलर, एक HTTP रिक्वेस्ट मैनेजर, और एक DOM पार्सर और एक्सट्रैक्टर के रूप में कार्य करता है।

rchipka/node-osmosis की मुख्य विशेषताएं क्या हैं?

rchipka/node-osmosis की मुख्य विशेषताएं हैं: Web Crawlers, Web Scraping and Extraction, Web Scraping Frameworks, DOM Tree Construction, CSS and XPath Query Engines, HTML Selector Extractors, CSS Selector, HTTP Request Management।

rchipka/node-osmosis के कुछ ओपन-सोर्स विकल्प क्या हैं?

rchipka/node-osmosis के ओपन-सोर्स विकल्पों में शामिल हैं: bda-research/node-crawler — node-crawler is a programmable web crawler for Node.js that manages request queues and automates data extraction. It… remitchell/python-scraping — This project is a Python web scraping library and automated data collection suite. It provides tools for extracting… lapwinglabs/x-ray — X-Ray is a web scraping framework and asynchronous web crawler designed to extract structured data from websites. It… symfony/dom-crawler — This project is an HTML and XML DOM parser designed for loading and navigating the structure of web documents to… ionicabizau/scrape-it — scrape-it is a Node.js web scraper and HTML parser designed to extract structured data from websites and HTML files.… ruipgil/scraperjs — Scraperjs is a JavaScript web scraping library and headless browser automation tool designed to extract structured…