Sha256: ff930cbcced2a1b80e29ca653b6edac02f29948164871a29a726dd392a6302fa

Contents?: true

Size: 956 Bytes

Versions: 7

Compression:

Stored size: 956 Bytes

Contents

= Anemone

Anemone is a web spider framework that can spider a domain and collect useful
information about the pages it visits. It is versatile, allowing you to
write your own specialized spider tasks quickly and easily.

See http://anemone.rubyforge.org for more information.

== Features
* Multi-threaded design for high performance
* Tracks 301 HTTP redirects to understand a page's aliases
* Built-in BFS algorithm for determining page depth
* Allows exclusion of URLs based on regular expressions
* Choose the links to follow on each page with focus_crawl()
* HTTPS support
* Records response time for each page
* CLI program can list all pages in a domain, calculate page depths, and more
* Obey robots.txt
* In-memory or persistent storage of pages during crawl, using TokyoCabinet or PStore

== Examples
See the scripts under the <tt>lib/anemone/cli</tt> directory for examples of several useful Anemone tasks.

== Requirements
* nokogiri
* robots

Version data entries

7 entries across 7 versions & 2 rubygems

Version Path
spk-anemone-0.4.0 README.rdoc
anemone-0.4.0 README.rdoc
anemone-0.3.2 README.rdoc
spk-anemone-0.3.1 README.rdoc
anemone-0.3.1 README.rdoc
spk-anemone-0.3.0 README.rdoc
anemone-0.3.0 README.rdoc