As I write this, the domain filing hasn’t come through, so I can only reach the site through the host’s third-level domain for now.
I set up a robots.txt forbidding every search engine from crawling any page on my site. But when I checked the stats today, there in the spider log was Tencent SOSO! I went and searched my site on SOSO, and there it was — indexed, at that absurdly long third-level domain, two pages a day.
When I tried to contact SOSO, I found there was no way to reach them at all. No contact information, and no complaint option on the cached snapshot either.
Every other search engine complies; only SOSO ignores the robots protocol. Screenshot evidence below:
Later I checked my robots file and found I’d written robot.txt instead of robots.txt. But every other search engine still didn’t crawl it — only SOSO. Is SOSO that dumb? The issue is fixed now, and I’ve blocked the search engines one by one.


