List Crawling: Guides to Crawlers, Lists, and Updates

About us

No post found

smart post preloader

No post found

smart post preloader

What crawling and a crawler actually are

Crawling is the act of visiting web pages on purpose and recording what you find. A crawler is the program that does the visiting. Search engines use crawlers to build an index of the public web. Smaller crawlers answer a narrower question: what is on this catalog, job board, directory, or URL list, and what changed since yesterday. This site is the guide to that work. It is not a dump of tool ads, and it is not a manual for opening pages you were not meant to open.

Crawl, crawler, spider, fetch, and scrape

Start with the words, because they get mixed. A crawl is one run. A crawler is the software that runs it. A spider is the same idea under an older name. Fetching is one request for one URL. Parsing is reading fields out of the page that came back. Scraping usually means that parsing step, or the whole job of turning pages into a table. Crawling, strictly, is discovery and fetching. On this blog we use list crawling for the common practical job: stay on pages that repeat the same item, follow that list to the end, and save one record per item.

Why a list page is an index, not the record

A list page is an index, not the record. A category grid, a job board, a city directory, or a search-results page stacks cards that share a shape: title, price, place, rating, and a link. The first guide is how to treat that shape as the unit of work. You name the entity, the fields, and the population, then you exhaust the index. Phase one collects every item URL once. Phase two fetches each URL and reads the full record. If the site already publishes an API or a CSV, that file is the better source.

Listing extraction and URL-list audits

Readers also meet a second meaning. In SEO and migration work, a list crawl means the opposite of discovery. You already have the URLs, from a sitemap, a redirect map, or a content inventory, and the crawler fetches that set and nothing else. After a move you are confirming a list, not finding one. Guides here label which job they mean.

Guides that stay, updates that are dated

The blog is organized as guides plus updates. A guide is a standing explanation: what a crawler does, how pagination continues, how to tell a finished list from a page-one file, how robots.txt differs from a contract. An update is dated. Vendors change defaults. Browsers change what a plain fetch can see. Sites move lists behind scripts. A court or a standard clarifies what a robots file is. Those notes say what changed, who it affects, and whether an older guide still holds.

What a careful crawler reads before it runs

What a careful crawler respects is part of the teaching, not a footer. RFC 9309 describes robots.txt as a request to automated clients, not as authorization. Read it, match your user-agent, and refresh it. Terms, copyright, and privacy law are separate questions. In the United States, hiQ Labs v. LinkedIn held that the Computer Fraud and Abuse Act does not, by itself, cover reading public pages that need no login. That ruling is not a permission slip. In the European Union, names, emails, and phone numbers are personal data even on a public page. This blog will not teach login bypasses or private contact harvesting.

Crawler, site, policy, and legal updates

Updates worth watching fall into four bins. Tooling: a crawler adds script rendering or changes its rate. Site structure: numbered pages become load-more, or a cursor replaces an offset. Policy: a host disallows a path, or publishes a feed you should use instead. Law: a robots rule is clarified, or a case draws a line between public pages and access you were not given. A short update should leave one decision: keep the job, change the fetcher, or stop.

Which guide to read before the next job

Use the site in this order. Read the definition guide if crawler is still fuzzy. Read the list-crawl guide before you design a job. Read the pagination note if a file came back short. Read the limits note before you store a column you did not mean to collect. Then check updates for the tool you are about to touch. You should leave knowing what crawling is, what a crawler is for, and how a list crawl differs from a full-site wander. Collect the index completely, read each record only once, and stop when the publisher has said no. This is a guide, not legal advice.

Scroll to Top