My blog has moved! Redirecting...

You should be automatically redirected. If not, visit http://www.dataminingblog.com and update your bookmarks.

Data Mining Research - dataminingblog.com: data mining books

I'm a Data Miner Collection (T-shirts, Mugs & Mousepads)

All benefits are given to a charity association.
Showing posts with label data mining books. Show all posts
Showing posts with label data mining books. Show all posts

Thursday, November 20, 2008

Data Mining Book: Know It All

Soon to be released (November, 21st), a new book about data mining: Data Mining: Know It All. The list of authors is impressive:

Soumen Chakrabarti, Earl Cox, Eibe Frank, Ralf Hartmut Güting, Jiawei Han, Xia Jiang, Micheline Kamber, Sam S. Lightstone, Thomas P. Nadeau, Richard E Neapolitan, Dorian Pyle, Mamdouh Refaat, Markus Schneider, Toby J. Teorey, Ian H. Witten.

Here is a description from Amazon:

This book brings all of the elements of data mining together in a single volume, saving the reader the time and expense of making multiple purchases. It consolidates both introductory and advanced topics, thereby covering the gamut of data mining and machine learning tactics ? from data integration and pre-processing, to fundamental algorithms, to optimization techniques and web mining methodology.

The proposed book expertly combines the finest data mining material from the Morgan Kaufmann portfolio. Individual chapters are derived from a select group of MK books authored by the best and brightest in the field. These chapters are combined into one comprehensive volume in a way that allows it to be used as a reference work for those interested in new and developing aspects of data mining.

This book represents a quick and efficient way to unite valuable content from leading data mining experts, thereby creating a definitive, one-stop-shopping opportunity for customers to receive the information they would otherwise need to round up from separate sources.

Continue reading... Sphere: Related Content

Thursday, October 23, 2008

Data Mining using SAS Enterprise Miner

I have recently found two new books about data mining. The author of these two books is Randall Matignon. He works at Amgen, Inc. in South San Francisco, California. He is a SAS/Microsoft Office VBA programmer with more than twenty years of experience. His expertise domains include pharmaceutical healthcare and biotechnology industries. Below is a short description of his two books.

Data Mining Using SAS Enterprise Miner

Data Mining Using SAS Enterprise Miner introduces the reader to a wide variety of data mining techniques in SAS® Enterprise Miner. This first-of-a-kind book explains the purpose of -- and reasoning behind -- every node that is a part of SAS® Enterprise Miner with regard to SEMMA design and SAS data mining analysis. Each chapter starts with a short introduction to the assortment of statistics that are generated from the various SAS® Enterprise Miner nodes, followed by detailed explanations of the configuration settings and the generated results that are located within each node. The end result of the author’s meticulous presentation is a well crafted study guide on the various methods that one employs to randomly sample, partition, transform, and filter the data within the process flow of SAS® Enterprise Miner. The book will explain the wide assortment of modeling designs that are available in addition to the process of assessing the various models under comparison in SAS® Enterpris e Miner v4.3.

Neural Network Modeling using SAS Enterprise Miner

Neural Network Modeling using SAS Enterprise Miner introduces the readers to a non-linear modeling design called neural network modeling using SAS Enterprise Miner. The book will also familiarize the readers with this predictive and classification methodology in statistics called neural network modeling. This book is designed in making statisticians, researchers and programmers aware of the awesome new product now available in SAS® called Enterprise Miner. This first of its kind book will reveal the strength and ease of use of the powerful new module in SAS® with step-by-step instructions in20creating a process flow diagram in preparation to data mining analysis and neural network predictive and classification modeling using SAS® Enterprise Miner v4.3.

For more information, visit www.sasenterpriseminer.com

Continue reading... Sphere: Related Content

Friday, February 08, 2008

Small book review: Super Crunchers

As written in an earlier post, Super Crunchers is a new book about data mining by Ian Ayres. Super crunching, according to Ayres is the action of applying data mining algorithms to real situations in order to make better decisions from data. I will make it clear right now: Super Crunchers will not give you examples of complex data mining techniques in real situations. Most of the book shows the use of randomized experiments (there are also a few pages on neural network, but that's all).

This book is nevertheless a very interesting reading for many reasons. First, the author did a very good job in introducing the basic ideas behind data mining for non-specialist readers. In addition, Ayres has collected a bunch of small, and very interesting, stories about people crunching data (wine quality prediction, baseball, etc.). In every situations, the author shows how crunching numbers help people make decisions but also how difficult it is to make non-expert believe in your results. This is, to my opinion, the most interesting aspect of the book.

Super Crunchers is very well written (I'm realizing now that I write that for most of my book reviews, but believe me you'll enjoy reading this book). After giving some examples, Ayres describes the actors of this industry (super crunching). He then introduces the idea of randomized experiments. There is also a nice chapter about the confrontation "Experts Versus Equations". He concludes by explaining why this enthusiasm for super crunching is happening only now and not before.

Finally, although the action of super crunching is certainly more about applying statistic methods (rather than data mining) to real situations, this is a must-have book, even for specialists in data mining. For interested reader, an interview of Ian Ayres is accessible here.

Ayres, I., Super Crunchers, Why Thinking-by-Numbers Is the New Way to be Smart, Bantam Books, 2007, 260p.

Continue reading... Sphere: Related Content

Friday, December 07, 2007

Super Crunchers

I have recently bought the book Super Crunchers: Why Thinking-by- Numbers Is the New Way to Be Smart by Ian Ayres. The book gives examples of data mining applications in nowadays companies. Instead of focusing on equations and algorithms, Ayres gives insights into data mining and statistics through real case studies. I will write a review about this book on Data Mining Research when I finish it.

If you're already interested, you can see the website of the book and a few words on the book on Newsweek.

[End of post]


Continue reading... Sphere: Related Content

Friday, November 16, 2007

Small book review: Web Dragons

Data mining is a field which is closely related to information extraction and search engines. Web Dragons: Inside the Myths of Search Engine Technology explains everything you want to know about search engines (the so called "web dragons") and how they work. Before reading the book, you perhaps wonder why Witten and co-authors called search engines "web dragons". After reading the book, I'm sure you will understand why. Search engines are guardians of the world information and their power is formidable.

The approach is descriptive and historical rather than technical. Thus, the book is intended to a wide audience: people working with data, librarians, webmasters, but also search engine users who wants to know more about the tool they use everyday. The first author, Ian Witten, is involved in the data mining field (see for example the famous book Data Mining (Witten and Frank, 2005). The book thus makes many allusions to data mining applications. It is divided as follows:

  • Setting the scene
  • Literature and the web
  • Meet the web
  • How to search
  • The web wars
  • Who controls information?
  • The dragons evolve
The two first chapters cover the history of search engines (starting from the very beginning: writing, etc.). You can easily skip these chapters (which maybe interesting to librarians for example) and start with the third one. There, you learn everything about the web, protocols, programming languages, etc. The strength of the book is to cover all these topics in a readable manner. You never face code or pseudo-code, only clear and interesting descriptions. The next chapter covers basics of search engine ranking (e.g. PageRank) in details and much more. Principal search engines are also introduced and explained. The following chapter (The web wars) explains the different ways of abusing such search engines (link boosting, term boosting, link farm, spam, etc.). The chapter is very interesting and instructing.

The next chapter (Who controls information?) points out the power of web dragons. They control world information and this raises privacy and copyright issues. Finally, the last chapter covers evolution of search engines. According to the authors, we are at the very beginning of information search. They focus on web communities that maybe the next step for search engine. As a conclusion, I recommend this book to anyone that is interested in how search engines work and especially how important they are for our society.

Web Dragons: Inside the Myths of Search Engine Technology (Witten et al., 2007).

Continue reading... Sphere: Related Content

Monday, March 19, 2007

Small book review: Java Data Mining

Unlike usual books on data mining discussed in this blog, Java Data Mining is a book written for data mining practitioners. Even if the word Java appears in the title, practitioners of other languages or software may be interested by the first part of the book (Strategy), which is really worth reading. The other parts of the book focus on the JDM API itself (Standards), problem solving with case study (Practice) and finally evolution of standards in data mining (Wrapping Up).

Data mining is clearly defined and compared with other concepts such as OLAP. A very interesting comparison is made between data mining and gold mining. As written previously, this book is practitioner-oriented. Moreover, the focus is on data mining with customer related information. Another good thing is the data mining glossary at the end of the book which is welcomed. According to the book, automated data mining strategies are being developed at KXEN. A related discussion can be found on Data Mining Research.

References to Wikipedia are to my opinion not appropriate in a book. Recent improvements have made this encyclopedia more reliable. However, I would never use it as a serious reference since anybody can write articles on it. A strange choice has been made regarding the word data (singular instead of plural). These small details aren't significant in regard to the quality of the book. To conclude, data mining practitioners and people using Java for data mining should really consider this book.

Continue reading... Sphere: Related Content

Monday, January 22, 2007

Stealing data mining books

Google is great in so many aspects. You can look for nearly anything and find relevant information in less than a second (I should advertise for Google :-) The dark side of such a powerful tool? It is often referencing illegal content. You were all aware of the possibility of downloading mp3, divx and so on. But did you know about data mining books?

If like me, you cannot believe it, then just check to see it with your eyes. Google pointed out this website as having a content related to data mining (in fact it is not wrong). Before I give you the link, I want to clearly state that accessing documents from it is illegal. I give this link as information and I believe that readers of my blog are responsible enough not to access these data (here is the link).

I don't know for you, but until today I wasn't aware that you can steal data mining books on the web...

Continue reading... Sphere: Related Content
 
Clicky Web Analytics