Scraper site

A scraper site is a website that copies content from other websites using web scraping. The purpose of creating such a site can be to collect advertising revenue or to manipulate search engine rankings by linking to other sites to improve their search engine ranking.

In the last few years scraper sites have proliferated at a high rate for spamming search engines. Open content is a common source of material for scraper sites.

A search engine is not a scraper site itself; sites such as Yahoo and Google gather content from other websites and index it so that the index can be searched with keywords. Search engines then display snippets of the original site content in response to a user's search.

Made for advertising

Some scraper sites are created to make money by using advertising programs. In such case, they are called Made for AdSense sites or MFA. This derogatory term refers to websites that have no redeeming value except to lure visitors to the website for the sole purpose of clicking on advertisements.[1]

Made for AdSense sites are considered sites that are spamming search engines and diluting the search results by providing surfers with less-than-satisfactory search results. The scraped content is considered redundant by the public to that which would be shown by the search engine under normal circumstances, had no MFA website been found in the listings.

Legality

Scraper sites may violate copyright law. Even taking content from an open content site can be a copyright violation, if done in a way which does not respect the license. For instance, the GNU Free Documentation License (GFDL)[2] and Creative Commons ShareAlike (CC-BY-SA)[3] licenses, used on Wikipedia,[4] require that a republisher inform readers of the license conditions, and give credit to the original author.

Techniques

Depending upon the objective of a scraper, the methods in which websites are targeted differ. For example, sites with mass amounts of content such as airlines, consumer electronics, department stores, etc. may be routinely targeted by their competition often to stay abreast of pricing information. Sophisticated scraping activity can be camouflaged by utilizing multiple IP addresses and timing search actions so they don't proceed at robot-like speeds and instead are more human like.

Some scrapers will pull snippets and text from websites that rank high for keywords they have targeted. This way they hope to rank highly in the search engine results pages (SERPs). RSS feeds are vulnerable to scrapers.

Some scraper sites consist of advertisements and paragraphs of words randomly selected from a dictionary. Often a visitor will click on a pay-per-click advertisement because it is the only comprehensible text on the page. Operators of these scraper sites gain financially from these clicks. Advertising networks claim to be constantly working to remove these sites from their programs, although there is an active polemic about this since these networks benefit directly from the clicks generated at this kind of site. From the advertisers' point of view, the networks don't seem to be making enough effort to stop this problem.

Scrapers tend to be associated with link farms and are sometimes perceived as the same thing, when multiple scrapers link to the same target site. A frequent target victim site might be accused of link-farm participation, due to the artificial pattern of incoming links to a victim website, linked from multiple scraper sites.

Domain hijacking

Main article: Domain hijacking

Some spammers who create scraper sites may hijack a recently expired domain name. Doing so will allow spammers to utilize the already-established search rankings for the domain name and incoming links. Some spammers may even try to match the topic of the expired site, to utilize their search rankings for those keywords. For example, an expired website for a photographer may be hijacked by a spammer who would generate a scraper site about photography tips.

See also

References

  1. Made for AdSense
  2. "Text of the GNU Free Documentation License".
  3. "Creative Commons Attribution-ShareAlike 3.0 Unported License".
  4. "Reusing Wikipedia Content".