Norconex Crawler

One open-source engine, two purpose-built variants — Web Crawler and File System Crawler — for collecting content from anywhere and turning it into clean, structured data ready for search, AI, or agents. Free, open source, and built to run at any scale.

Norconex Crawler V4 is here, in beta — a rebuilt core for horizontal scalability and clustering, plus a new visual configurator. Meet the Crawler · Try the Visual Configurator

Looking for a production-ready release today? V3 is the current stable version.

One Engine, Two Variants

Besides the fetchers and a few source-specific pieces, both variants share the same crawl → process → commit core. Pick the one that matches where your content lives.

Web Crawler

Best for websites, portals, and authenticated web experiences.

  • Crawl modern JavaScript-heavy sites with full browser rendering
  • Handle authenticated areas and logged-in user journeys
  • Respect sitemaps, redirects, canonical URLs, and site rules
  • Tune requests with custom headers, cookies, and user agents
  • Skip unchanged content on repeat crawls

Web Quick Start

File System Crawler

Best for drives, shares, repositories, and internal file estates.

  • Crawl local disks, network shares, and mounted repositories
  • Connect to SharePoint, Nextcloud, object storage, and HDFS
  • Reach remote content over FTP, SFTP, and WebDAV
  • Traverse large file estates with path and depth controls
  • Capture content and metadata consistently across every source

File System Quick Start

Features

Universal Output

Send crawled content to search engines, databases, queues, and custom targets without locking into one destination.

Built to Extend

Adapt crawler behavior in plain Java when the out-of-the-box pipeline isn't enough.

Embeddable or Standalone

Run it from the command line or embed the same engine inside your own Java applications.

Metadata & Content Control

Enrich, filter, rename, and transform content and metadata before anything is sent downstream.

Portable Operations

Reuse configurations, isolate environment-specific settings, and promote the same crawl setup across environments.

Resumable & Observable

Resume after interruptions, tap into crawler events, and diagnose failures with detailed logs.

See All Features

Need Help At Scale?

Our team offers configuration, deployment, and support packages for teams running the crawler in production — from a single-site setup to enterprise deployments across multiple sources. Contact us for a quote.

Get Started

The crawler is free and open source — download the stable release and run your first crawl today, or preview what's coming in V4. Need a hand at scale? We're here.