Twiceler - What?!

jeko

Lounge-Member
Hallo zusammen,

viele von euch kennen den Bot Twiceler wahrscheinlich schon... Also meine Logs quellen fast über von seinen Anfragen, heute waren's wieder mal etwas gegen die 3'500.

Wenn man sich auf der schlichten Homepage des Unternehmens umschaut, kommt bei mir schnell der Eindruck eines kleinen Hinterstübchenbüros auf. Viel Informationen gibt die Page ja nicht her, aber das Team scheint sehr qualifiziert zu sein...

Weiss da jemand mehr...? Also was für eine Motivation steckt dahinter, einfach mehr Informationen. Find's irgendwie "beängstigend" wenn so ein Unternehmen so anonym auftritt. Zudem hält er sich auch nicht, wie versprochen, an die robots.txt. Hab ihn seit 2 Wochen ausgesperrt, doch trotzdem besucht er die Seite alle 2-3 Tage mal und dann gleich richtig...

Grüsse
jeko
 
Hmm... Ich weiss eben nicht was ich von ihm halten soll. Aussperren wär weniger das Problem. Würde nur gerne wissen, wer/was das ist, falls jemand von euch was weiss :icon7:
Hab ihnen mal ein Mail geschrieben (wie "sie" es ja selbst wünschen), mal gespannt was zurückkommt. Im Wesentlichen halt, dass Twiceler ziemlich häufig vorbeikommt und meine robots.txt nicht beachtet.
 
@ Jeko

Zitat von Cuill:
"Webmaster Information
Twiceler is an experimental robot. The user-agent is “twiceler.” It could take 24-48 hours for us to re-read your robots.txt file. If you need something blocked immediately, please let us know."

Es funktioniert!

Gruß
Harry
 
Danke harrygrey :)

Ich hab die ganze Angelegenheit etwas aus dem Blick verloren, aber nochmal kurz in meine Logs geschaut und siehe da, twiceler hielt sich zurück. Scheint also doch zu funktionieren.

Hab hier noch das Mail, dass ich geschrieben hab. Zum Antworten kam ich leider nicht mehr, da wie gesagt, andere Sachen um die Ohren gehabt...
Dear Dominique,

Twiceler is an experimental crawler that we are developing for our new search engine.
It is important to us that it obey robots.txt, and that it not crawl sites that do not wish to be
crawled. It would help us debug the crawler if you could send us some examples of its failing
to obey your robots.txt.

If you wish I will glad to add xxx.ch to our list of sites to exclude, and
I apologize for any inconvenience this has caused you.

Please feel free to contact me if you have any further questions.

Sincerely,

James Akers
Operations Engineer
Cuill, Inc.


>
>> *From: *Dominique Sandoz <jeko@gmx.ch <mailto:jeko@gmx.ch>>
>> *Date: *August 17, 2007 12:56:16 PM PDT
>> *To: *contact@cuill.com <mailto:contact@cuill.com>
>> *Subject: **Twiceler is a little bit too eagerly*
>>
>> Dear Sir/Madam,
>>
>> As written on your site (Cuill - Twiceler Information), I assumed Twiceler would obey robots.txt. I changed my robots.txt (http://xxx.ch/robots.txt) weeks ago, that my file xxx.html (http://xxx.ch/xxx.html) isn't indexed anymore by your bot. Well, it's still doing. Even today I got over 3200 requests from your bot, starting at my list.html and following all the links on this page.
>>
>> Maybe Twiceler will revolutionate searching in the net. But please, could you make Twiceler following robots.txt?
>>
>> I'm open for critics too, tell me if I've done something wrong in my robots.txt.
>>
>> Thanks for your attention,
>> Dominique Sandoz
Yap, my English is...
 
Na, ich folge ja streng dem Motto "don't be evil"... Oh Mist, verletz' ich grad'n Copyright...?

Wie gesagt, die Frequenz ist 'runter und ich hab eigentlich 'n Flair für Hinterstübchenprojekte...:)
 
Zurück
Oben