花了好一些時間研究 robots.txt 的寫法與設定

對於劣質的網路爬虫類(spider , BOT)實在防不勝防,想全部擋掉,但又想在 google , yahoo , MSN live 上被人找到

在未研究 robots.txt 之前,實在不是件簡單的事

導入主題~有哪些劣質的 spider 呢?google一下,發現已有 外國網友 整理

(只要複製、另存在WWW 根目錄 / robots.txt 即可跟不少 bad spiders say Goodbye 了!)

 

網路上有不少 robots.txt checker ,可以檢測我們寫的 robots.txt 語法是否正確,並給我們一些建議

紅使用 New Robots.txt Syntax Checker 來進行檢測

以紅為例子,在它的 input field 輸入  http://red.ns2go.com/robots.txt    即可

會逐行檢查語法,並依各個 BOT 是否能 通過(Allow) / 封鎖(Disallow) 將其標示顏色

1.有錯誤處,會標示紅色(檢查時,即發現兩、三個 BOT 重覆定義,刪除即可 )

2.有其他建議,在最下方亦會告知

3.額外加入從後台發現的不知名 robots ,加以限定

 

當完全通過檢測時,網站會給予一個小貼貼,將 code 複製後貼在網站上,可提供檢核

snap044.gif

======================================================

設定說明:僅允許Googlebot , MSNBot , Slurp(Yahoo bot) , BaiDuSpider 來訪

間隔來訪時間:3600 sec (1小時) 紅覺得 bot 不用這麼常來吧,1小時來一次已足夠

其他列表的 bot 皆不允許抓取資料

倒數 10 行指的是:不允許任何 BOT 抓取 wp-admin , wp-contend …等目錄中的資料

並且指定 sitemap 檔的路徑為 http://red.ns2go.com/sitemap.xml.gz

=================以下為紅的 robots.txt====================== 

User-agent: Googlebot
Crawl-delay: 3600
Disallow:

User-agent: MSNBot
Crawl-delay: 3600
Disallow:

User-agent: Slurp
Crawl-delay: 3600
Disallow:

User-agent: BaiDuSpider
Crawl-delay: 3600
Disallow:

User-agent: Gecko
Disallow: /

User-agent: YodaoBot
Disallow: /

User-agent: BotRightHere
Disallow: /

User-agent: WebZip
Disallow: /

User-agent: larbin
Disallow: /

User-agent: b2w/0.1
Disallow: /

User-agent: Copernic
Disallow: /

User-agent: psbot
Disallow: /

User-agent: Python-urllib
Disallow: /

User-agent: URL_Spider_Pro
Disallow: /

User-agent: CherryPicker
Disallow: /

User-agent: EmailCollector
Disallow: /

User-agent: EmailSiphon
Disallow: /

User-agent: WebBandit
Disallow: /

User-agent: EmailWolf
Disallow: /

User-agent: CopyRightCheck
Disallow: /

User-agent: Crescent
Disallow: /

User-agent: SiteSnagger
Disallow: /

User-agent: ProWebWalker
Disallow: /

User-agent: CheeseBot
Disallow: /

User-agent: LNSpiderguy
Disallow: /

User-agent: Alexibot
Disallow: /

User-agent: Teleport
Disallow: /

User-agent: TeleportPro
Disallow: /

User-agent: MIIxpc
Disallow: /

User-agent: Telesoft
Disallow: /

User-agent: Website Quester
Disallow: /

User-agent: moget/2.1
Disallow: /

User-agent: WebZip/4.0
Disallow: /

User-agent: WebStripper
Disallow: /

User-agent: WebSauger
Disallow: /

User-agent: WebCopier
Disallow: /

User-agent: NetAnts
Disallow: /

User-agent: Mister PiX
Disallow: /

User-agent: WebAuto
Disallow: /

User-agent: TheNomad
Disallow: /

User-agent: WWW-Collector-E
Disallow: /

User-agent: RMA
Disallow: /

User-agent: libWeb/clsHTTP
Disallow: /

User-agent: asterias
Disallow: /

User-agent: httplib
Disallow: /

User-agent: turingos
Disallow: /

User-agent: spanner
Disallow: /

User-agent: InfoNaviRobot
Disallow: /

User-agent: Harvest/1.5
Disallow: /

User-agent: Bullseye/1.0
Disallow: /

User-agent: Mozilla/4.0 (compatible; BullsEye; Windows 95)
Disallow: /

User-agent: Crescent Internet ToolPak HTTP OLE Control v.1.0
Disallow: /

User-agent: CherryPickerSE/1.0
Disallow: /

User-agent: CherryPickerElite/1.0
Disallow: /

User-agent: WebBandit/3.50
Disallow: /

User-agent: NICErsPRO
Disallow: /

User-agent: Microsoft URL Control – 5.01.4511
Disallow: /

User-agent: DittoSpyder
Disallow: /

User-agent: Foobot
Disallow: /

User-agent: SpankBot
Disallow: /

User-agent: BotALot
Disallow: /

User-agent: lwp-trivial/1.34
Disallow: /

User-agent: lwp-trivial
Disallow: /

User-agent: BunnySlippers
Disallow: /

User-agent: Microsoft URL Control – 6.00.8169
Disallow: /

User-agent: URLy Warning
Disallow: /

User-agent: Wget/1.6
Disallow: /

User-agent: Wget/1.5.3
Disallow: /

User-agent: Wget
Disallow: /

User-agent: LinkWalker
Disallow: /

User-agent: cosmos
Disallow: /

User-agent: moget
Disallow: /

User-agent: hloader
Disallow: /

User-agent: humanlinks
Disallow: /

User-agent: LinkextractorPro
Disallow: /

User-agent: Offline Explorer
Disallow: /

User-agent: Mata Hari
Disallow: /

User-agent: LexiBot
Disallow: /

User-agent: Web Image Collector
Disallow: /

User-agent: The Intraformant
Disallow: /

User-agent: True_Robot/1.0
Disallow: /

User-agent: True_Robot
Disallow: /

User-agent: BlowFish/1.0
Disallow: /

User-agent: JennyBot
Disallow: /

User-agent: MIIxpc/4.2
Disallow: /

User-agent: BuiltBotTough
Disallow: /

User-agent: ProPowerBot/2.14
Disallow: /

User-agent: BackDoorBot/1.0
Disallow: /

User-agent: toCrawl/UrlDispatcher
Disallow: /

User-agent: WebEnhancer
Disallow: /

User-agent: suzuran
Disallow: /

User-agent: TightTwatBot
Disallow: /

User-agent: VCI WebViewer VCI WebViewer Win32
Disallow: /

User-agent: VCI
Disallow: /

User-agent: Szukacz/1.4
Disallow: /

User-agent: QueryN Metasearch
Disallow: /

User-agent: Openfind data gatherer
Disallow: /

User-agent: Openfind
Disallow: /

User-agent: Xenu's Link Sleuth 1.1c
Disallow: /

User-agent: Xenu's
Disallow: /

User-agent: Zeus
Disallow: /

User-agent: RepoMonkey Bait & Tackle/v1.01
Disallow: /

User-agent: RepoMonkey
Disallow: /

User-agent: Microsoft URL Control
Disallow: /

User-agent: Openbot
Disallow: /

User-agent: URL Control
Disallow: /

User-agent: Zeus Link Scout
Disallow: /

User-agent: Zeus 32297 Webster Pro V2.9 Win32
Disallow: /

User-agent: Webster Pro
Disallow: /

User-agent: EroCrawler
Disallow: /

User-agent: LinkScan/8.1a Unix
Disallow: /

User-agent: Keyword Density/0.9
Disallow: /

User-agent: Kenjin Spider
Disallow: /

User-agent: Iron33/1.0.2
Disallow: /

User-agent: Bookmark search tool
Disallow: /

User-agent: GetRight/4.2
Disallow: /

User-agent: FairAd Client
Disallow: /

User-agent: Gaisbot
Disallow: /

User-agent: Aqua_Products
Disallow: /

User-agent: Radiation Retriever 1.1
Disallow: /

User-agent: Flaming AttackBot
Disallow: /

User-agent: Oracle Ultra Search
Disallow: /

User-agent: MSIECrawler
Disallow: /

User-agent: PerMan
Disallow: /

User-agent: searchpreview
Disallow: /

User-agent: TurnitinBot
Disallow: /

User-agent: wget
Disallow: /

User-agent: ExtractorPro
Disallow: /

User-agent: WebZIP/4.21
Disallow: /

User-agent: WebZIP/5.0
Disallow: /

User-agent: HTTrack 3.0
Disallow: /

User-agent: TurnitinBot/1.5
Disallow: /

User-agent: WebCopier v3.2a
Disallow: /

User-agent: WebCapture 2.0
Disallow: /

User-agent: WebCopier v.2.2
Disallow: /

User-agent: *       
Sitemap: http://red.ns2go.com/sitemap.xml.gz
Crawl-delay: 3600 
Disallow: /wp-admin/
Disallow: /wp-content/
Disallow: /wp-includes/
Disallow: /wp-login.php
Disallow: /awstats/
Disallow: /awstats-icon/ 
Disallow: /temp/       

 

參考資料:

◇.robots at 2007-07-09  Kirin Lin 我上網、改程式、部落格,故我在-改版中-

◆.SEO tools:Robot Control Code Generation Tool  robot.txt 產生器,可以用用看

◇.New Robots.txt Syntax Checker robots.txt checker

相關文章:

[Wedding] 文喬Joe & 家玉Joyce《迎取》
[Wedding] 文喬Joe & 家玉Joyce《迎取》

一起看夜景的幸福
一起看夜景的幸福

[Wedding] 文喬Joe & 家玉Joyce《宴客》
[Wedding] 文喬Joe & 家玉Joyce《宴客》

小莫Part II @微笑的秘密花園
小莫Part II @微笑的秘密花園

回家
回家

[Betrothal] 伊成 & 欣慧《迎取》
[Betrothal] 伊成 & 欣慧《迎取》

Tags: