Webbots, Spiders, and Screen Scrapers and over one million other books are available for Amazon Kindle. Learn more

Sell Us Your Item
For a $1.73 Gift Card
Trade in
Have one to sell? Sell yours here
Start reading Webbots, Spiders, and Screen Scrapers on your Kindle in under a minute.

Don't have a Kindle? Get your Kindle here, or download a FREE Kindle Reading App.
Sorry, this item is not available in
Image not available for
Color:
Image not available

To view this video download Flash Player

 

Webbots, Spiders, and Screen Scrapers: A Guide to Developing Internet Agents with PHP/CURL [Paperback]

Michael Schrenk
4.6 out of 5 stars  See all reviews (19 customer reviews)


Available from these sellers.


Free Two-Day Shipping for College Students with Amazon Student

Formats

Amazon Price New from Used from
Kindle Edition $17.57  
Paperback --  
Sell Back Your Copy for $1.73
No matter where you bought them, get up to 70% back when you sell your books at Amazon.com.
Used Price$10.00
Trade-in Price$1.73
Price after
Trade-in
$8.27
There is a newer edition of this item:
Webbots, Spiders, and Screen Scrapers: A Guide to Developing Internet Agents with PHP/CURL Webbots, Spiders, and Screen Scrapers: A Guide to Developing Internet Agents with PHP/CURL 3.5 out of 5 stars (17)
$23.03
In stock on May 25, 2013

Book Description

March 30, 2007 1593271204 978-1593271206 First Edition,Annotated

The Internet is bigger and better than what a mere browser allows. Webbots, Spiders, and Screen Scrapers is for programmers and businesspeople who want to take full advantage of the vast resources available on the Web. There's no reason to let browsers limit your online experience-especially when you can easily automate online tasks to suit your individual needs.

Learn how to write webbots and spiders that do all this and more:

Programmatically download entire websites Effectively parse data from web pages Manage cookies Decode encrypted files Automate form submissions Send and receive email Send SMS alerts to your cell phone Unlock password-protected websites Automatically bid in online auctions Exchange data with FTP and NNTP servers

Sample projects using standard code libraries reinforce these new skills. You'll learn how to create your own webbots and spiders that track online prices, aggregate different data sources into a single web page, and archive the online data you just can't live without. You'll learn inside information from an experienced webbot developer on how and when to write stealthy webbots that mimic human behavior, tips for developing fault-tolerant designs, and various methods for launching and scheduling webbots. You'll also get advice on how to write webbots and spiders that respect website owner property rights, plus techniques for shielding websites from unwanted robots.

As a bonus, visit the author's website to test your webbots on sample target pages, and to download the scripts and code libraries used in the book.

Some tasks are just too tedious-or too important!- to leave to humans. Once you've automated your online life, you'll never let a browser limit the way you use the Internet again.



Editorial Reviews

About the Author

Michael Schrenk uses webbots and data-driven web applications to create competitive advantages for businesses. He has written for Computerworld and Web Techniques magazines and has taught courses on Web usability and Internet marketing. He has also given presentations on intelligent Web agents and online corporate intelligence at the DEFCON hacker's convention.


Product Details

  • Paperback: 328 pages
  • Publisher: No Starch Press; First Edition,Annotated edition (March 30, 2007)
  • Language: English
  • ISBN-10: 1593271204
  • ISBN-13: 978-1593271206
  • Product Dimensions: 7.3 x 1 x 9.1 inches
  • Shipping Weight: 1.4 pounds
  • Average Customer Review: 4.6 out of 5 stars  See all reviews (19 customer reviews)
  • Amazon Best Sellers Rank: #114,153 in Books (See Top 100 in Books)

More About the Author

Michael Schrenk is a software developer, author and instructor. He specializes in automated web browsing agents known as webbots. His book, "Webbots, Spiders, & Screen Scrapers" (2007, No Starch Press, San Francisco) is the definitive work on the subject.

Michael uses the Internet in new and innovative (odd?) ways to provide competitive advantages for his clients in The US, Europe and Asia.

He also helps journalists more effectively use computers to conduct online research through automation and by describing where and how to find otherwise hidden online information. No stranger to Europe--he's lived in Moscow and Madrid, Mike taught at the 2008 European Investigative Journalism Conference (Brussels Belgium), twice in 2009 he lectured at The Center for Investigative Journalism (London England) and later in 2009, he lead several sessions at the VVOJ Journalism conference (Utrecht The Netherlands).

Last August, Mike made his fourth speaking appearance at the DEFCON computer hacking conference. Mike lives in sunny Las Vegas, Nevada (USA). You can contact him at http://www.schrenk.com

Watch this book's video at:
http://www.youtube.com/watch?v=B7z7tP74RxQ

Customer Reviews

Most Helpful Customer Reviews
37 of 40 people found the following review helpful
2.0 out of 5 stars Does the basics. December 5, 2007
By Brian
Format:Paperback|Amazon Verified Purchase
"Webbots, Spiders, adn Screen Scrapers" is a solid book for building basic scripts to do web scraping. Michael Schrenk goes covers the "should you do this" aspect very well, and devotes much of the book to these kinds of topics. On that reason alone I give him major kudos, "just because you CAN do a thing, doesn't mean you SHOULD."

Technically the book and examples are very basic and beginner level. All code is procedural and has absolutely no references to object oriented programming at all. This is great for a simple project, but building anything larger than a targetted webbot or two is beyond the scope of this book.

I was very dismayed at Mr. Schrenk's opinion of regular expressions:
"The use of regular expressions is a parsing language in itself, and most modern programming languages support aspects of regular expressions. In the right hands, regular expressions are also useful for parsing and substituting text; however, they are famous for thier sharp learning curve and cryptic syntax. I avoid regular expressions whenever possible."

This disregard for regular expressions effectively wipes out a powerful toolset for budding developers. Regular expressions are no harder to learn than PHP. The reasons for his disdain for them is also flawed:

"The regular expression engine used by PHP is not as efficient as engines used in other languages, and is certainly less efficient than PHP's built-in functions for parsing HTML."

PHP uses the same regular expression engine used (very effectively) in PERL with the use of the preg_* functions. There has been many studies that show preg_* style expressions outperform basic text matching in PHP. In this assesment the author is terribly wrong.

The book does a great job of explaining how to make single use scripts for scraping, but never how to create a larger infrastructure. There is no focus on creating multi process engines with pcntl_fork(), or proc_open(), these are critical for scaling web scraping applications. A single script scraping a few hundred websites on a single thread would take ages over a multi-threaded engine.

If you are looking to break into web scraping and not sure where to start, this is likely the best (and possibly only) book on the market. If you are intermediate or advanced you will quickly question the author's logic and see that scaling will become the number one issue you have to over come.
Was this review helpful to you?
24 of 26 people found the following review helpful
3.0 out of 5 stars Solid introduction to webbots, with a catch. April 27, 2007
Format:Paperback
I picked up this book full of enthusiasm, spiders are just plain cool, they go out and start downloading data for you, reading webpages, and even understanding them a little. My enthusiasm was dashed a little however on page four: “You may use any of the scripts in this book for your own personal use, as long as you agree not to redistribute them... and agree not to sell or create derivative products under any circumstances.”. I develop in PHP professionally, and a lot of the code I write ends up getting used somewhere with some sort of a for-profit basis, which pretty effectively prevents me from using any code between the covers (at its strictest reading, I’m not sure I can even change the code).

The book does a great job of introducing different sorts of web agents that you can create programatically (more than just spiders) and introduces all sorts of interesting projects along those lines. Throughout the book a series of libraries written by the author are leveraged to make the retrieval and parsing of the various pages much easier. While newer developers will enjoy being able to concentrate on the “big picture” I found myself itching for more information on the nitty gritty.

Some of the projects explored include: price monitoring, image capturing (want to be your own google image search? :) ), link verification, spiders, and snipers. Each of the different projects received it’s own chapter, and effectively covered a lot of the topics covered within.

Overall, I would recommend this book to beginner to intermediate PHP developers looking to tackle the world of web agents, it’s a good primer on the related topics, and at the very least will give you some ideas on the complexities involved. As their skill grows they will probably find them-self either moving past the libraries included with the book, or modifying them greatly. My biggest complaint is the lack of coverage on the robots.txt file, some talk is given to it in terms of blocking robots from your own site, but I didn’t see any code that actually dealt with parsing it for your own robot.
Was this review helpful to you?
10 of 10 people found the following review helpful
5.0 out of 5 stars Great Book with Lots of Information August 25, 2007
Format:Paperback|Amazon Verified Purchase
This book covers every aspect I could ever hope a book on web bots would cover. It goes into great detail and provides lots of background information about things such as why you should use web bots, security issues, how to authenticate a bot with password protected sites, writing search engine crawlers, parsing HTML, how to handle cookies, HTTP headers, dealing with forms and a lot more.

I was very pleased with how this book covered concepts. The book uses PHP and the cURL library as a teaching tool instead of trying to give a lesson in how to use PHP as a crawler language. The way the code is explained makes it very easy to translate into whatever language you are most comfortable coding in. The book uses fundamental functional programming concepts which make it easy to pick up the general idea without actually knowing PHP.

My boss bought this book to help my group us with a project we were working on, and even my co-workers who had no background with PHP were able to use this book to write a web bot in C# (using the cURL library) very easily. The concepts from this book easily transfered over to object-oriented concepts.
Comment | 
Was this review helpful to you?
Most Recent Customer Reviews
5.0 out of 5 stars A Good Explination of WebBot Technology
As a long time developer I have often found my self having to write up screen scrapers and crawlers to go out and automate various processes across the net. Read more
Published 9 months ago by Daniel A. Lewis
5.0 out of 5 stars Automating data collection with your eyes closed...
This is a review of Michael's 2nd Edition of the same book (I received an early release copy from the publisher, I did not have an opportunity to read the 1st edition):... Read more
Published 14 months ago by Gregory Zentkovich
3.0 out of 5 stars If it's all you want to do...
I was hoping for more from this book due to the reviews but it's not all that. As pointed out, it's has solid basics. Read more
Published 18 months ago by Dark Wing Duck
4.0 out of 5 stars Great book!
If you want to 'automate' your browsing then this is a great book, with examples for every conceivable application. Read more
Published on August 31, 2010 by Don Rogers
5.0 out of 5 stars Best for this subject
The power of this book is not so much in it's code examples but rather in it's ability to change your perspective. Read more
Published on February 12, 2010 by Dwight R. Schofield
5.0 out of 5 stars Great introduction
This is a great introduction on the subject. The supplied PHP library does all the work.
Published on October 8, 2009 by introfini
5.0 out of 5 stars This book is useful
This book is not like very algorithmic, but you can know the basic of webbots writing and some techniques involved. Read more
Published on January 25, 2009 by Ching C. Nang
5.0 out of 5 stars Great Basic Book
Need to learn how to browse the web with your own software instead of manually browsing? The is the best book on the subject. Read more
Published on December 2, 2008 by Joe Todd
5.0 out of 5 stars a super introduction to web spiders
I won't re-iterate the excellent reviews already posted on this book, other than to say this is probably my favorite all-time programming book: excellently written, highly... Read more
Published on November 16, 2008 by Yannick Pouliot
5.0 out of 5 stars :-) bots
This book is a great reference and/or introduction to the cURL library. After reading this book, I realized it is not intended as a single solution for bot programming. Read more
Published on August 6, 2008 by C. D. Cox
Search Customer Reviews
Only search this product's reviews




Forums

Search Customer Discussions
Search all Amazon discussions

Start a new discussion
Topic:
First post:
Prompts for sign-in
 




So You'd Like to...


Create a guide


Look for Similar Items by Category