花了好一些時間研究 robots.txt 的寫法與設定
對於劣質的網路爬虫類(spider , BOT)實在防不勝防,想全部擋掉,但又想在 google , yahoo , MSN live 上被人找到
在未研究 robots.txt 之前,實在不是件簡單的事
導入主題~有哪些劣質的 spider 呢?google一下,發現已有 外國網友 整理
(只要複製、另存在WWW 根目錄 / robots.txt 即可跟不少 bad spiders say Goodbye 了!)
網路上有不少 robots.txt checker ,可以檢測我們寫的 robots.txt 語法是否正確,並給我們一些建議
紅使用 New Robots.txt Syntax Checker 來進行檢測
以紅為例子,在它的 input field 輸入 http://red.ns2go.com/robots.txt 即可
會逐行檢查語法,並依各個 BOT 是否能 通過(Allow) / 封鎖(Disallow) 將其標示顏色
1.有錯誤處,會標示紅色(檢查時,即發現兩、三個 BOT 重覆定義,刪除即可 )
2.有其他建議,在最下方亦會告知
3.額外加入從後台發現的不知名 robots ,加以限定
當完全通過檢測時,網站會給予一個小貼貼,將 code 複製後貼在網站上,可提供檢核
======================================================
設定說明:僅允許Googlebot , MSNBot , Slurp(Yahoo bot) , BaiDuSpider 來訪
間隔來訪時間:3600 sec (1小時) 紅覺得 bot 不用這麼常來吧,1小時來一次已足夠
其他列表的 bot 皆不允許抓取資料
倒數 10 行指的是:不允許任何 BOT 抓取 wp-admin , wp-contend …等目錄中的資料
並且指定 sitemap 檔的路徑為 http://red.ns2go.com/sitemap.xml.gz
=================以下為紅的 robots.txt======================
User-agent: Googlebot
Crawl-delay: 3600
Disallow:
User-agent: MSNBot
Crawl-delay: 3600
Disallow:
User-agent: Slurp
Crawl-delay: 3600
Disallow:
User-agent: BaiDuSpider
Crawl-delay: 3600
Disallow:
User-agent: Gecko
Disallow: /
User-agent: YodaoBot
Disallow: /
User-agent: BotRightHere
Disallow: /
User-agent: WebZip
Disallow: /
User-agent: larbin
Disallow: /
User-agent: b2w/0.1
Disallow: /
User-agent: Copernic
Disallow: /
User-agent: psbot
Disallow: /
User-agent: Python-urllib
Disallow: /
User-agent: URL_Spider_Pro
Disallow: /
User-agent: CherryPicker
Disallow: /
User-agent: EmailCollector
Disallow: /
User-agent: EmailSiphon
Disallow: /
User-agent: WebBandit
Disallow: /
User-agent: EmailWolf
Disallow: /
User-agent: CopyRightCheck
Disallow: /
User-agent: Crescent
Disallow: /
User-agent: SiteSnagger
Disallow: /
User-agent: ProWebWalker
Disallow: /
User-agent: CheeseBot
Disallow: /
User-agent: LNSpiderguy
Disallow: /
User-agent: Alexibot
Disallow: /
User-agent: Teleport
Disallow: /
User-agent: TeleportPro
Disallow: /
User-agent: MIIxpc
Disallow: /
User-agent: Telesoft
Disallow: /
User-agent: Website Quester
Disallow: /
User-agent: moget/2.1
Disallow: /
User-agent: WebZip/4.0
Disallow: /
User-agent: WebStripper
Disallow: /
User-agent: WebSauger
Disallow: /
User-agent: WebCopier
Disallow: /
User-agent: NetAnts
Disallow: /
User-agent: Mister PiX
Disallow: /
User-agent: WebAuto
Disallow: /
User-agent: TheNomad
Disallow: /
User-agent: WWW-Collector-E
Disallow: /
User-agent: RMA
Disallow: /
User-agent: libWeb/clsHTTP
Disallow: /
User-agent: asterias
Disallow: /
User-agent: httplib
Disallow: /
User-agent: turingos
Disallow: /
User-agent: spanner
Disallow: /
User-agent: InfoNaviRobot
Disallow: /
User-agent: Harvest/1.5
Disallow: /
User-agent: Bullseye/1.0
Disallow: /
User-agent: Mozilla/4.0 (compatible; BullsEye; Windows 95)
Disallow: /
User-agent: Crescent Internet ToolPak HTTP OLE Control v.1.0
Disallow: /
User-agent: CherryPickerSE/1.0
Disallow: /
User-agent: CherryPickerElite/1.0
Disallow: /
User-agent: WebBandit/3.50
Disallow: /
User-agent: NICErsPRO
Disallow: /
User-agent: Microsoft URL Control – 5.01.4511
Disallow: /
User-agent: DittoSpyder
Disallow: /
User-agent: Foobot
Disallow: /
User-agent: SpankBot
Disallow: /
User-agent: BotALot
Disallow: /
User-agent: lwp-trivial/1.34
Disallow: /
User-agent: lwp-trivial
Disallow: /
User-agent: BunnySlippers
Disallow: /
User-agent: Microsoft URL Control – 6.00.8169
Disallow: /
User-agent: URLy Warning
Disallow: /
User-agent: Wget/1.6
Disallow: /
User-agent: Wget/1.5.3
Disallow: /
User-agent: Wget
Disallow: /
User-agent: LinkWalker
Disallow: /
User-agent: cosmos
Disallow: /
User-agent: moget
Disallow: /
User-agent: hloader
Disallow: /
User-agent: humanlinks
Disallow: /
User-agent: LinkextractorPro
Disallow: /
User-agent: Offline Explorer
Disallow: /
User-agent: Mata Hari
Disallow: /
User-agent: LexiBot
Disallow: /
User-agent: Web Image Collector
Disallow: /
User-agent: The Intraformant
Disallow: /
User-agent: True_Robot/1.0
Disallow: /
User-agent: True_Robot
Disallow: /
User-agent: BlowFish/1.0
Disallow: /
User-agent: JennyBot
Disallow: /
User-agent: MIIxpc/4.2
Disallow: /
User-agent: BuiltBotTough
Disallow: /
User-agent: ProPowerBot/2.14
Disallow: /
User-agent: BackDoorBot/1.0
Disallow: /
User-agent: toCrawl/UrlDispatcher
Disallow: /
User-agent: WebEnhancer
Disallow: /
User-agent: suzuran
Disallow: /
User-agent: TightTwatBot
Disallow: /
User-agent: VCI WebViewer VCI WebViewer Win32
Disallow: /
User-agent: VCI
Disallow: /
User-agent: Szukacz/1.4
Disallow: /
User-agent: QueryN Metasearch
Disallow: /
User-agent: Openfind data gatherer
Disallow: /
User-agent: Openfind
Disallow: /
User-agent: Xenu's Link Sleuth 1.1c
Disallow: /
User-agent: Xenu's
Disallow: /
User-agent: Zeus
Disallow: /
User-agent: RepoMonkey Bait & Tackle/v1.01
Disallow: /
User-agent: RepoMonkey
Disallow: /
User-agent: Microsoft URL Control
Disallow: /
User-agent: Openbot
Disallow: /
User-agent: URL Control
Disallow: /
User-agent: Zeus Link Scout
Disallow: /
User-agent: Zeus 32297 Webster Pro V2.9 Win32
Disallow: /
User-agent: Webster Pro
Disallow: /
User-agent: EroCrawler
Disallow: /
User-agent: LinkScan/8.1a Unix
Disallow: /
User-agent: Keyword Density/0.9
Disallow: /
User-agent: Kenjin Spider
Disallow: /
User-agent: Iron33/1.0.2
Disallow: /
User-agent: Bookmark search tool
Disallow: /
User-agent: GetRight/4.2
Disallow: /
User-agent: FairAd Client
Disallow: /
User-agent: Gaisbot
Disallow: /
User-agent: Aqua_Products
Disallow: /
User-agent: Radiation Retriever 1.1
Disallow: /
User-agent: Flaming AttackBot
Disallow: /
User-agent: Oracle Ultra Search
Disallow: /
User-agent: MSIECrawler
Disallow: /
User-agent: PerMan
Disallow: /
User-agent: searchpreview
Disallow: /
User-agent: TurnitinBot
Disallow: /
User-agent: wget
Disallow: /
User-agent: ExtractorPro
Disallow: /
User-agent: WebZIP/4.21
Disallow: /
User-agent: WebZIP/5.0
Disallow: /
User-agent: HTTrack 3.0
Disallow: /
User-agent: TurnitinBot/1.5
Disallow: /
User-agent: WebCopier v3.2a
Disallow: /
User-agent: WebCapture 2.0
Disallow: /
User-agent: WebCopier v.2.2
Disallow: /
User-agent: *
Sitemap: http://red.ns2go.com/sitemap.xml.gz
Crawl-delay: 3600
Disallow: /wp-admin/
Disallow: /wp-content/
Disallow: /wp-includes/
Disallow: /wp-login.php
Disallow: /awstats/
Disallow: /awstats-icon/
Disallow: /temp/
參考資料:
◇.robots at 2007-07-09 Kirin Lin 我上網、改程式、部落格,故我在-改版中-
◆.SEO tools:Robot Control Code Generation Tool robot.txt 產生器,可以用用看
◇.New Robots.txt Syntax Checker robots.txt checker
![[Wedding] 文喬Joe & 家玉Joyce《迎取》](http://lh3.ggpht.com/_2l8BzEPgrEk/TTW-t1OKmWI/AAAAAAAAFBA/SSd_aKS6uQc/s160-c/2011.01.01-236.jpg)

![[Wedding] 文喬Joe & 家玉Joyce《宴客》](http://lh6.ggpht.com/_2l8BzEPgrEk/TTXAStq25xI/AAAAAAAAFEU/4SPyq6fHcII/s160-c/2011.01.01-323.jpg)


![[Betrothal] 伊成 & 欣慧《迎取》](http://lh6.ggpht.com/_2l8BzEPgrEk/TJYm-fNaXZI/AAAAAAAAEPY/cQJAoy8RtA8/s160-c/2010.09.12-63.jpg)
1. Comment by Red
14/八月/2007 at 12:19 下午
從設定 robots.txt 當天起,至今有兩天的時間
有發現三個改變:
1.從後端的點擊數(hits)可以發現少了約一半
2.且很多莫名的瀏覽器也都不見了
3.符合 robots.txt 中的機器人的來訪次數也少了許多,不再是一天就來找尋數百次,伺服器都可能被爬掛了 =.=
robots 終於不再這麼煩人了~讚!
2. Comment by SIKO
24/十月/2007 at 12:55 上午
感謝啦,
那我就不客氣的接收下來了
By SIKO