Tencent SOSO Ignores the Robots Protocol

As I write this, the domain filing hasn’t come through, so I can only reach the site through the host’s third-level domain for now.

I set up a robots.txt forbidding every search engine from crawling any page on my site. But when I checked the stats today, there in the spider log was Tencent SOSO! I went and searched my site on SOSO, and there it was — indexed, at that absurdly long third-level domain, two pages a day.

When I tried to contact SOSO, I found there was no way to reach them at all. No contact information, and no complaint option on the cached snapshot either.

Every other search engine complies; only SOSO ignores the robots protocol. Screenshot evidence below:

QQ20130521191323

QQ20130521191730

Later I checked my robots file and found I’d written robot.txt instead of robots.txt. But every other search engine still didn’t crawl it — only SOSO. Is SOSO that dumb? The issue is fixed now, and I’ve blocked the search engines one by one.