My blog has moved! Redirecting...

You should be automatically redirected. If not, visit http://www.dataminingblog.com and update your bookmarks.

Data Mining Research - dataminingblog.com: vulgarization

I'm a Data Miner Collection (T-shirts, Mugs & Mousepads)

All benefits are given to a charity association.
Showing posts with label vulgarization. Show all posts
Showing posts with label vulgarization. Show all posts

Thursday, February 14, 2008

The numbers behind Numb3rs

Do you watch Numb3rs? It is a TV show where the two main characters are a FBI agent (Don) and his younger brother (Charlie), a very clever mathematician. The show is based on the idea that crimes can be solved with numbers. In each episode, Charlie uses different mathematic techniques to find information regarding some crime scene. In an episode of the third season (Brutus), Charlie uses data mining techniques.

To be honest, I haven't watched this episode. I'm not a big fan of the show, but I find it interesting sometimes. If, like me, you missed this episode, Devlin and Lorden have written a book about the show: The Numbers Behind Numb3rs: Solving Crime with Mathematics (Plume Book, 2007).

This book describes the mathematics behind Numb3rs. Chapter 3 (on data mining) is quite interesting. Authors are mentioning facts and anecdotes relevant to data mining. However, it is often not related to the show itself (it goes deeper in explaining the details of some methods). The book tries to vulgarize the concepts of data mining. It is quite normal since the audience certainly consists of people interested by the show.

The problem is that sometimes, the text is so vulgarized that it is wrong. Look at the quote below, taken from the page 28 of the book:

"Neural networks: special kinds of computer programs that can predict the probability of crimes and terrorist attacks."
Oops! This is not vulgarization anymore. It is simply wrong. This is the bigger risk of vulgarization: trying to change the terms to make an explanation clearer may ends up incorrect. However, the rest of the chapter is generally fine and even sometimes very technical.

Continue reading... Sphere: Related Content

Wednesday, February 06, 2008

Stupid Data Miner Tricks

I have recently read an interesting article regarding data mining entitled Stupid Data Miner Tricks: Overfitting the S&P 500. Indeed, the paper is written in a somehow provocative manner (as can be seen in the title). Since the paper is old (1995), you may already have heard about it.

It is written by David J. Leinweber who was a PhD at Caltech. In his article Leinweber mentions the bad connotation of the expression data mining, a few decades ago. He then warns of the many dangers of data mining when it is badly used. By the way, if you want examples of reliable use of data mining nowadays (I mean outside universities), take a look at Super Crunchers (I will soon post a review of this interesting book).

As written by Leinweber, his paper gives an example of "[...] totally bogus application of data mining in finance." For this, he shows the strong statistical correlation between the annual changes of the S&P 500 stock index and the butter production in Bangladesh (!) The focus of the paper is thus the problem of overfitting, which happens when a model fits to well a training set. The model has thus bad generalization abilities when evaluated on the test set. I will conclude by citing a sentence from the article, which is, to my opinion, always an issue in data mining: "When doing this kind of analysis [regression] it's important to be very careful what you ask for, because you'll get it."

Continue reading... Sphere: Related Content

Friday, October 19, 2007

WIRED point of view on AI

In its October 2007 issue, WIRED has special section named "Geekipedia". In this supplement, WIRED summarizes 149 people, facts or concepts that they think are important. Among the list, one can find "Artificial Intelligence". The description is quite negative and focus on different aims that AI hasn't been able to achieve. I agree with them on the first half of the explanation regarding AI. They write that "[...] while researchers have built awesome technology, they've failed to grapple with philosophy". AI researchers can tell me if I'm wrong, but I think this sentence can be considered as acceptable. However, in the middle of the text, things start to go wrong.

Although the message WIRED intends to give about AI (i.e. AI hasn't yet achieved most of its initial aims), they give bad examples. They write that "[...] computers failed one commonsense task after another [...]". The problem doesn't come from this sentence, rather from examples of such "tasks".

First example: computer fails to understand natural languages. Oops! Very bad example. Although one can discuss the meaning of the word "understand", it is clear that speech processing, recognition and synthesis are examples of successful applications in machine learning. The second example they give is even worst. They write that a computer cannot distinguish a dog from a cat. Oops again! Face recognition is one of the best example of machine learning success story. And there is only one step from the Human to the animal. Indeed, I have a colleague in machine learning who is doing face recognition on a cat database... and it's working!

But the worst is yet to come (yes, believe me). The last paragraph explaining AI contains the following sentence: "Nowadays, Google "knows" pretty much anything you ask it. But its insanely fast and powerful work is modestly described as data-mining, not thinking". Out of the spelling, I'm surprised by the bad connotation given to the data mining term. So, although their description is not completely wrong, they haven't chosen the best examples to illustrate the limitations of AI. There are still a lot of things a computer cannot do, so examples are not missing...

Continue reading... Sphere: Related Content
 
Clicky Web Analytics