awesome-repositories.com
Blog
MCP
awesome-repositories.com

Discover the best open-source repositories with AI-powered search.

ExploreCurated searchesOpen-source alternativesSelf-hosted softwareBlogSitemap
ProjectMCP serverAboutHow we rankPress
LegalPrivacyTerms
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
joeyism avatar

joeyism/linkedin_scraper

0
View on GitHub↗
3,746 stars·880 forks·Python·gpl-3.0·16 views

Linkedin Scraper

This project is a LinkedIn data scraper and professional profile extractor designed to collect information from professional networking sites. It functions as a headless browser scraper that extracts professional profiles, company details, and job listings using automated browser sessions.

The tool includes a session manager that saves and loads authentication cookies to maintain persistent access to protected profiles. It employs configurable browser settings and user-agent mimicry to simulate human activity and bypass bot detection.

Data extraction capabilities cover person profiles, company overviews, social feed posts, and job listings filtered by keywords and location. The system also supports the retrieval of contact details, education, and work experience.

Features

  • Headless Browser Automation - Uses a headless browser to simulate human interaction and render dynamic content for data extraction.
  • Headless Browser Orchestrators - Implements a headless browser system to simulate human interaction and extract dynamic professional profile data.
  • Job Market Data - Extracts job listings and requirements to facilitate analysis of hiring trends and industry demands.
  • Job Market Scraping - Collects job requirements and application links using keyword and location filters.
  • Professional Profile Scraping - Gathers career data and company history from professional networking sites using automated browser tools.
  • DOM-Based Extractions - Extracts professional data by querying the HTML structure of rendered web pages using JavaScript.
  • Profile Extractors - Gathers work experience, education, and contact details from individual professional user pages.
  • Session Management - Manages authentication cookies and browser settings to access protected professional profiles.
  • Session and Credential Management - Captures and stores user session tokens and credentials in files for persistence across multiple requests.
  • Browser Session Authentication - Utilizes browser cookies to authenticate automated requests to protected professional profiles.
  • Session-Based Scrapers - Utilizes browser session cookies to scrape restricted profiles, company details, and job listings from LinkedIn.
  • Session-Cookie Persistences - Saves and reuses authentication cookies in local files to maintain persistent access across script runs.
  • Web Session Management - Coordinates the login process and maintains state to access restricted areas of the professional network.
  • Company Intelligence Utilities - Gathers business overviews and industry classifications to build professional company profiles.
  • Organizational Data Extraction - Collects industry details, organizational size, and headquarters locations from company information pages.
  • Competitive Market Research - Provides capabilities to extract company overviews and industry data for competitive analysis.
  • Professional Networking Automation - Automates the collection of posts and profile updates to monitor industry activity.
  • Contact Discovery - Identifies and extracts contact information associated with professional user profiles.
  • Post & Comment Scraping - Fetches post content, timestamps, and engagement metrics from organization social feeds.
  • Lead Enrichment - Extracts and augments lead profiles with contact information and career history from professional networking sites.
  • Profile Information Retrieval - Fetches specific personal and professional profile details using unique identifiers.
  • User Profile Retrieval - Retrieves detailed work experience, education, and skill sets from professional user profiles.
  • User Agent Rotation - Cycles through user agent strings and viewport settings to mimic human behavior and avoid bot detection.
  • Search Filtering Logic - Narrows down job listings and professional profiles based on user-provided keywords and location criteria.
  • Browser Argument Configuration - Configures headless browser arguments, viewport size, and user agents to simulate human activity.

Star history

Star history chart for joeyism/linkedin_scraperStar history chart for joeyism/linkedin_scraper

How this analysis was created: This summary and feature list are AI-generated from collected project material and can contain mistakes. Stars, license and language are imported from GitHub. Inclusion does not mean that we have tested or audited this project. Check the source documentation for any feature you depend on. Learn more on our About page.

AI search

Explore more awesome repositories

Describe what you need in plain English — the AI ranks thousands of curated open-source projects by relevance.

Start searching with AI

Projects sharing features with Linkedin Scraper

These projects share indexed features with Linkedin Scraper. Shared tags can include platform or build tooling; verify the primary use case before treating a result as a replacement.
  • xchaoinfo/fuck-loginxchaoinfo avatar

    xchaoinfo/fuck-login

    5,871View on GitHub↗

    This project is a web scraping authentication tool designed to automate login simulations and bypass authentication walls. It provides utilities for programmatically navigating login screens to gain access to protected web content and restricted data. The software specifically enables session establishment through QR code authentication, allowing users to scan a code with a mobile device to capture and save authentication cookies. These captured sessions are used to maintain persistence for future automated requests. The tool covers broader capabilities in browser-based automation, including

    Pythonloginpython3weibo
    View on GitHub↗5,871
  • guyungy/damaihelperGuyungy avatar

    Guyungy/damaihelper

    2,551View on GitHub↗

    Damaihelper is a ticketing automation bot and browser automation framework designed to monitor ticket availability and execute checkout processes. It utilizes a ticket purchasing script to automate the selection and purchase of tickets on web platforms based on predefined user criteria. The tool includes a graphical user interface for managing scripts and configuring automation parameters, allowing users to trigger tasks without using a command line. To maintain access, it employs browser session management to save and reuse authentication cookies, avoiding repetitive manual login procedures.

    HTML
    View on GitHub↗2,551
  • garrytan/gstackgarrytan avatar

    garrytan/gstack

    110,596View on GitHub↗

    gstack is an AI agent framework and development workflow system designed to automate the software development lifecycle. It coordinates specialized AI personas to manage tasks across product design, engineering management, and quality assurance, transforming product intent into technical specifications and final releases. The project is distinguished by its deep integration of headless browser automation and semantic code memory. It utilizes a persistent Chromium daemon for web scraping and visual auditing, and implements a searchable knowledge base that logs architectural decisions and repos

    TypeScript
    View on GitHub↗110,596
  • projectdiscovery/katanaprojectdiscovery avatar

    projectdiscovery/katana

    15,584View on GitHub↗

    Katana is a web crawler and spider designed for security reconnaissance and web application mapping. It functions as a utility for identifying endpoints, forms, and API structures across web targets by combining standard HTTP request traversal with headless browser automation to render dynamic, JavaScript-heavy content. The tool distinguishes itself through its ability to maintain authenticated sessions and handle complex web interactions, such as automated form submission and captcha resolution. It provides granular control over the discovery process, allowing users to define specific crawl

    Goclicrawlergocrawler
    View on GitHub↗15,584
Compare all 30 related projects→

Frequently asked questions

What does joeyism/linkedin_scraper do?

This project is a LinkedIn data scraper and professional profile extractor designed to collect information from professional networking sites. It functions as a headless browser scraper that extracts professional profiles, company details, and job listings using automated browser sessions.

What are the main features of joeyism/linkedin_scraper?

The main features of joeyism/linkedin_scraper are: Headless Browser Automation, Headless Browser Orchestrators, Job Market Data, Job Market Scraping, Professional Profile Scraping, DOM-Based Extractions, Profile Extractors, Session Management.

Which projects share features with joeyism/linkedin_scraper?

Projects with overlapping indexed features include: guyungy/damaihelper — Damaihelper is a ticketing automation bot and browser automation framework designed to monitor ticket availability and… xchaoinfo/fuck-login — This project is a web scraping authentication tool designed to automate login simulations and bypass authentication… garrytan/gstack — gstack is an AI agent framework and development workflow system designed to automate the software development… projectdiscovery/katana — Katana is a web crawler and spider designed for security reconnaissance and web application mapping. It functions as a… huaying/instagram-crawler — This project is a web scraping and automation tool designed to collect public data from Instagram and perform… joeanamier/tiktokdownloader — TikTokDownloader is a containerized automation tool designed for the systematic collection and archiving of social…