It's not a very difficult problem to solve. 95% of the content-farm spam comes from a few domains. In the same way that spam-blacklists have proved to be the most-effective way to combat e-mail spam, Google just needs to decide to shut these content-farms out. They don't need to do anything sophisticated like tweak their algorithm... just shut them out. The fact that it hasn't been done yet suggests to me that Google doesn't want to.
But if Google shuts those domains out, then the content farmers can just pick up new ones. Furthermore, Google would need to have some sort of process for determining which domains are content-farm spam before shutting them out.
As with spam, the bad guys can be far more agile than the good guys. The point of having algorithms is that once you have a good algorithm, your search-quality staff doesn’t have to scale linearly with the number of pages on the Web.
(Disclaimer: I work in the search division of Nokia, i.e., I work for one of Google’s competitors.)
Google has 23,000 employees and ~$10B in yearly operating income. I'm pretty sure they could stay ahead of the bad guys if they wanted to.
EDIT: Additionally, it is hard to game pagerank rapidly because the rest of the web needs to link to your site. So, even if you switch your knock-off wikipedia site to a new domain, it would take weeks/months to rise through the rankings. I'm pretty sure it would only take a few employees (at most) to stay ahead of these huge spam sites.
Compared to how many millions of people making money on the web, all of whom have an incentive to boost their own traffic regardless of whether that's the best thing for users?
Joy's Law comes into play here: "No matter who you are, most of the smartest people work for someone else." 23,000 employees sounds huge, until you compare it with the million+ people who make their living as full-time E-bay sellers, and the however many million people who make their living off AdSense.
Google has 23,000 employees ... . I'm pretty sure they could stay ahead of the bad guys if they wanted to.
In the past there were people within Google who had the skill, knowledge, organizational connections and authority to quickly and gracefully reduce the problem of link spam without harming the bottom line too much.
I think we all loved their early work, and our love helped propel their little company to multibillions-per-year.
I don't think that kind of unique early experience and subsequent problem-solving effectiveness is snap-in interchangeable.
Some of those "rock stars" work at Facebook now. Google might not be capable of fixing the problem gracefully without them.
it is hard to game pagerank rapidly because the rest of the web needs to link to your site
Not really. Some people own different websites, purchased through different accounts, hosted on different servers. They, then, link between those sites.
well, yeah, that's how the algorithms work... by using a training dataset where a human actually says "is spam."
And as for the "is spam" button... a given email isn't automatically added to the spam training dataset the minute 1 guy hits the spam button. it goes automatically into YOUR spam folder but not automatically into the training dataset. That takes many "votes" from many users.
That's why I would be happy with a personal spam filter: training by users' votes can be gamed, and producing false positives can take legitimate websites out of business.
I see there are some browser plugins floating around: I'm not happy with those because I use multiple workstations, and because hacking around a product's deficiencies is not "voting with my wallet".
In my experience, the path from “engineer recognizes that queries X, Y, and Z return crap search results” to “search engine improves its performance with queries X, Y, and Z without creating more crap somewhere else” is more difficult than a lot of people realize.
Still, a power law applies and you can get huge gains by punishing the biggest offenders and moving down from there. And yes, spammers can always restart, but the sandbox helps with that.
High-end Nokia phones have a local-search client that competes with Google Maps. The front end for people using regular browsers is here: http://maps.ovi.com/
It reminds me of the reddit IAMA a while back by an affiliate marketer [1]. There were lots of responses by people claiming he was the scum of the earth, poisoning the web, not adding value, but he just calmly stated his position that affiliate marketers are a significant source of income for Google, and "Google == value", therefore affiliate marketing has value.
Same might be said for content farms, so I reckon your last sentence is on the money.
Perhaps because of legal issues, particularly when Google is being sued for tweaking its ranking algorithm to boost their own pages higher.
I dont understand the legalities of that lawsuit though. A ranking is an opinion. Its not that its some fundamental physical constant or a formally well defined quantity. Can I be sued for my subjective judgment ?
OT: Will appreciate if you let me know why you down voted.
I don't understand those lawsuits either. Isn't Google a private company? Aren't they entitled to do what they want with their product? I see no reason why their results must be "fair", legally.
IANAL, but proving that a company has a monopoly and used that monopoly in an abusive way is very hard.
Consider that when discussing whether a company has a monopoly on a market, alternatives and cost of switching also comes up (besides market share); so it will be even harder than in the case of Microsoft.
Then even if Google is discovered to have a monopoly, shutting down a website ... can be argued that it was in the best interests of its users and faithful to the product's original mission: i.e. nobody can sue Microsoft for improving Windows / not bundling third-party software (they got sued for extending Windows with new functionality that destroyed competition in an existing market).
can be argued that it was in the best interests of its users and faithful to the product's original mission
If you're willing to pay the tens of millions of dollars and take the PR hit that is a government anti-trust lawsuit, you may eventually have the privilege of making that argument before a federal judge, who probably isn't technology literate enough to comprehend it. Let's face facts here, this conversation goes over the heads of 80% of internet users. The judge is going to see "the federal government accuses them of restraining this site's trade by removing them from the search index, and Google admits to it" and then that's the ballgame.
Or maybe the government never gets around to suing Google. Or maybe they get a judge who knows this sort of stuff already. But that's still a scenario keeping Google's Legal Dept. up at night.
Yeah, but that would be like suing Microsoft for making Windows more secure because that can "restrain the trade of companies producing Antiviruses".
That would be a pretty stupid argument, even in the face of a non-technical jury, wouldn't it?
Of course, the negative image would hurt, but lets be honest, the suit against Microsoft accomplished basically nothing it couldn't handle, while waisting taxpayer's money. Are these antitrust suits getting started so easily?
I would be surprised to see them suffering for optimising the algorithm to try and trace the original source of content though. With the sort of sites we're talking about, by definition that original source will be more up-to-date content which the user could reasonably expect to be presented preferentially. At which point - you're welcome to make your business from rehosting public content that originates elsewhere, but if you expect us to primarily direct people to your copy rather than the original then I think you're onto a loser.
Let's put it another way. I could conceivably build a valuable service by taking content from StackOverflow and Wikipedia (or wherever) and automatically linking the two to enable people to get some more context around some questions, maybe reformatting pages to enable side-by-side content or something similar. It wouldn't be a trivial service but it wouldn't be impossible in the least and it could plausibly add value. As such it wouldn't be unreasonable to preferentially direct users in some cases to that source rather than the original, as algorithmically optimised content - the preferential ranking would be a result of the value of the linking algorithm. Without this though, by what measure am I conceivably superior to the original source by having an out-of-date copy with fewer legitimate inbound links and more irrelevant content (adverts)?
Being a monopoly isn't illegal. Using your monopoly power in a way that might harm another company isn't illegal. What's illegal is doing that unreasonably and capriciously, as Microsoft were with Netscape in that trial, or DR and Lotus in previous trials.
No law, as far as I am aware, forbids a monopoly from producing a crappy product. Antitrust law comes into play when a competitor offers a less-crappy product and the monopoly tries to drive the competitor out of business.
No, but a company that has a monopoly cannot manipulate a market by making a product selectively crappy, like a search engine that fails to find its competitors or a program that fails to run on a competing, compatible OS (early win3 on DR-DOS).
Yes, if you make your product selectively crappy in order to undermine competition, it’s an antitrust problem. Microsoft making Windows not run on DR-DOS is an antitrust problem. Microsoft making Windows fail at multiuser security is not an antitrust problem.
2. Saying that it only comes from a few sources is dangerous thinking. This type of thinking was used by politicians to get a ban on adult services from Craigslist ("get rid of it there and it goes away!").
A good email blacklist (such as the Spamhaus Zen list) will detect 90% of spam with no false positives and negligible load.
There is a lot of evidence that there are not many botnets and each one has little diversity of control. This observation does not generalize to other kinds of spam such as 419s.