Saturday, May 30, 2009

Fresh answers and old ones

Sometime the answer to a question is changing according to time. In this case, Cerberus bougth Crysler for 7.4 Billions in 2007. Anyway, the company is currently evaluating a Chapter 11 and it sold a large part of his assets to Fiat, an Italian company.

Business is a circle. Fiat was in deep deep troubles 5 years ago. Then, everything changed when Sergio Marchione, the new manager accepted to run the company as CEO.

Ask is nailing the fresh answer, Google and Live give the old one, Yahoo is not nailing the story.

Estimating the ImpressionRank of Web Pages

ImpressionRank is the number of times users viewed the page, while browsing search results returned by the search engine. In turn, ImpressionRank has an intuitive correlation with power law search query distribution.

Estimating the ImpressionRank of Web Pages proposes a number of methodologies for estimating ImpressionRank. It is based on sampling of search suggestion services provided by search engines and on extraction of popular keywords extracted by web pages. The paper achieves quite interesting results in terms of recall and convergence speed.

I am pretty sure that the work is quite interesting for all the SEO companies.



Friday, May 29, 2009

Twitter Search is soaring

Search is alive and kicking

In the last couple of months we have seen many completely new initiatives.

Are we now out of the crisis? I say, YES we are!
  1. Microsoft Bing, a serious new competitor in the arena. In my opinion Bing is moving in a direction of more focused search. There are some ideas inspired by Vivisimo, Seachme , and my old Snaket for the categories. Some other ideas taken from the Ask3D interface. A much better ranking, and more important a new way of accessing and organizing search results. It seems that Microsoft wants to invest more money in search and is a very good news for the whole sector. I will follow their performance in the next months.
  2. Wolfram Alpha, is something very different and also very difficult to evaluate. They aims at having all systematic knowledge immediately computable by anyone. A very ambitious goal. For this reason they have my full respect. They provide amazing results for some queries and very poor results for other queries. Precision is very high, recall is low. I will observe their performance more and more in the next months.
  3. Many proposals for RealTime Search with players such as Twitter Search, OneRiot, FriendFeed, and the many others. This is where freshness is important.
Having different search options is an absolute need for a truly open access to the information. Information belongs to people, and search engines are the gateway to it. Therefore is very good to say that Search is alive and kicking.

Wednesday, May 27, 2009

Real Time Semantic Correlations

Algorithm is the key. Always. Here I am talking about fresh correlations among automatically extracted fresh semantic entities. This is a funny story, but I just wanted to explain a technology we had since many years.

Noemi Letizia is the new supposed italian Monica Lewinsky. Someone is saying that she has an affair with Silvio Berlusconi, this is not proved. But there are rumors. Veronica Lario is Berlusconi's wife. She asked for a divorce due to this story. Gino Flaminio is the ex-boyfriend of Noemi Letizia. He had an important interview with the main italian newspaper. I guess you know why Clinton is there.

Algorithm nailed the true essence of the story

Tuesday, May 26, 2009

Real time Semantic Search

Well everyone is talking about Semantic Search ( have you seen Wolfram Alpha ?). You know, I live in this place so far away from all the rest ;-) But here in this little and wasted land we keep asking a question ;-) !!
Now do you know the answer? Did you ever hear about it?



I think that Semantic Search needs Freshness and realtime data as well

Sunday, May 24, 2009

Google vs. the Real-Time Web

Google vs. the Real-Time Web is a nice article by the way of Gigaom.
  • "By contrast, real-time discovery engines like Twitter and Facebok use a more dynamic kind of democracy, linking to content that users finds worthwhile. As a result, content on the web is splitting into two basic models, and understanding this distinction makes clear why Google’s centralized role is being threatened"
I like the comments:
  • "There is no time available to develop meta data that separates the wheat from the chaff, or in recent terms, the stupid bacon jokes from real news about Swine Flu"
  • "It has to have a past to give people’s reactions time to develop. But yes, right now if you say realtime, you certainly can get funded. In the Valley, at least."
This is what I think:
  • I partially agree with Adam. We already saw some working Real time web search. Namely News blending into Web search. Now, Real time is not just news blending. It’s much more, but the news blending experience may be leveraged there.

Using Graphics Processors for High Performance IR Query Processing

Using Graphics Processors for High Performance IR Query Processing is a paper exploring the use of CUDA GPU as co-processor for serving search queries. The idea is quite intriguiging since GPUs may potentially offer amazing performances at very low cost.

The authors propose a parallel sum prefix based Rice encoding and a PForDelta encoding.
Rice coding encodes an integer (the gap between two consecutive docIDs) by choosing a number of bits b such that 2^b is close to the average of all the gaps, and then representing each integer
as q · 2^b + r for some r < 2b. Then the integer is encoded in a unary part, consisting of q 1s followed by a 0, and a binary part of b bits representing r. PForDelta first selects a value b such that most gaps are less than 2^b, and then uses an array of b-bit values to store all gaps less then 2^b while all other gaps are stored in a special format as exceptions.

The paper describes gap encoding, decoding, and merge operations. In addition, it discusses how to process ranked query, skip lists, dijunctive and conjiuntive queries.

Performances are good, a single server with a GPU and a CPU we can sustain a query arrival rate beyond 300 q/s, versus less than 100 for CPU only. The index size was ~25M documents.

Quite an interesting paper, indeed.

Saturday, May 23, 2009

Learning to Tag

Learning to Tag is a Microsoft paper about suggesting tags associated to Flicker images. The authors use traditional textual co-occurences and visual features (such as colour histogram, and moment). These features are combined by using a RankBoost learning process.

The results are compared with a naive linear combination of ranking signals, and with simple tag co-occurences. It seems that Microsoft likes to combine different ranking signals with RankBoost, since I saw many papers describing variations of this idea and the performance they achieve seems quite good.

Friday, May 22, 2009

American Idol: Freshness, Variety and UI

Who is the winner of american idol? Let's see what is the performance of various search engines. I cannot vote since I am involved in the contest. Anyway, these are the criteria:

1) Freshness, how fast do they broke the news
2) Accurateness, how much accurate are the results
3) Variety, do they provide a sense of variety in them (not just 10 blue links?)
4) UI, is the user experience good?

What is your opinion?


Ask.com has the winner on the top with News, Video and Audio



Google has the official site, wikipedia, and youtube videos



Yahoo has the news with images



Live has the news with images




Searchme has the best UI all around (IMO)



Twitter has all the rumors, and they broke the story





Hmm ... actually OneRiot broke the story !!

Mapping the World's Photos

Mapping the World's Photos is a fascinating paper from Jon Kleimberg et al. They used textual features and SIFT image signatures to geo-localize a large sample of Flicker images. It seems that the two different class of features show a mutual benefit.

It's fascinating to discover what are the most photographed world landmarks.. Applestore in 5th avenue, NYC is the 5th most photograped place in the world, while Rome's colosseum is just 41th...

Thursday, May 21, 2009

Mining Interesting Locations and Travel Sequences From GPS Trajectories

Mining Interesting Locations and Travel Sequences From GPS Trajectories . Hmm ... when I read this paper from Microsoft, I though "...we will follow you ". The key idea is to rank users and places using a HITS-based approach since there is an evident mutual reinforcement propriety among them. 107 people were tracked for one year.

Wednesday, May 20, 2009

Network Analysis of Collaboration Structure in Wikipedia

Network Analysis of Collaboration Structure in Wikipedia is a study about topic polarization and edit activities on Wikipedia. It seems that there are topics where people starts neverending revert wars. I wonder why the authors are not talking about spam and bots here....

Tuesday, May 19, 2009

Understanding User's Query Intent with Wikipedia

"Understanding User's Query Intent with Wikipedia" is a nice paper from Microsoft. It explains how to classify Web queries in order to trigger results from different verticals (such as news, images, video, travel, shopping).

Each category is bootstrapped with few keywords chosen by editors. These descriptions are then automatically expanded using a random walk on wikipedia's categories and concepts. The semantic concepts are then extracted by using Gabrilovich's esa. Results are provided on Live search query log for July 2007 (which has just ~2.6M frequent queries). Precision, Recall and F1 measures are quite impressive and this generic solution can compete with the best ad-hoc KDD2005 classification competition result.

Query classification is a very important topic for Search Engines, and leveraging Wikipedia is definitevely a good idea. (see also my previous posting for Yahoo's query classification)

Multidimension Scaling

Multidimension Scaling is a technique for projecting multi-dimensional points in a 2-d plan, so that the distances as preserved as much as possible.

Here you have a C++ code skeleton with STL and boost.

Monday, May 18, 2009

Range minimum queries and LCA

Range Minimum Query (RMQ) is used on arrays to find the position of an element with the minimum value between two specified indices. There is a trivial quadratic solution based on dynamic programming. I found this tutorial complete and useful. Many other solutions exist. In particular, I appreciated the one based on Segment Trees. A bookmark, and I plan to revisit it soon.