Decision Mechanics

Insight. Applied.

  • Services
    • Decision analysis
    • Big data analysis
    • Software development
  • Articles
  • Blog
  • Privacy
  • Hire us

Dirty data is the biggest challenge facing data scientists

February 11, 2015 By editor

A recent survery of data scientists by CrowdFlower found that, when it comes to challenges

Dirty data is the #1 hurdle…

My own experience leads me to agree. Access to good quality data remains a huge problem. Organizations would do well to invest in improving the quality of their data before boosting their analytics capabilities. Garbage in, garbage out.

One error I see made regularly is not to appreciate that it’s cheaper to fix problems “upstream”—e.g. at the point the data is collected. Better tools, UX and training can significantly improve the quality of the data entering your analysis ecosystem.

Too many organizations see cleaning as a process that occurs centrally, late in the collection pipeline. By then it’s often too late to fix the problem, and all you can do is discard the data. Even worse, you may not notice the problem and use the inaccurate data in your modeling.

Many errors can only be identified in context—and as the data moves further from its origin, context disappears. For instance, if a skate park employee enters the ages of a bunch of customers as being 80, a simple iPad app could ask for confirmation, based on statistical profiling.

However, if the same data were, instead, cleaned by an analyst at HQ a week later, it would be impossible to tell if that was an error or the reunion of the 1955 Olympians. And, discarding unusual, but accuate, data is going to reduce your ability to spot emerging trends, or niche markets.

Filed Under: General

What’s hot at the 2015 Strata+Hadoop World conference?

February 11, 2015 By editor

It’s always refreshing to see data scientists turn the tools on themselves. Gives the field credibility.

Benedikt Koehle looked at the frequency of ngrams in the abstracts for the 2015 Strata+Hadoop World conference. He concluded, as a result of this analysis, that

2015 will be probably known as the “Spark Strata”

He also notes the resurgence of interest in R, at the expense of Python.

I also notice that the word “APIs” is very prominent—paralleling the growth of interest in other development-related areas.

Filed Under: Big data

Paying by Numbers—Should Data Scientists Receive Performance Based Compensation?

February 8, 2015 By editor

Learning Tree International just published another of my articles. I recently read an opinion that data analysts will start to be paid for performance. My article is presents an opposing view. While there are jobs where performance can be measured in isolation, I don’t believe that data science is one of them.

Read the article for my full argument.

Filed Under: Data analysis Tagged With: invited article

Free on-line data science courses from Stanford

February 7, 2015 By editor

Stanford has a number of free on-line courses in session that might be of interest to data scientists. They started last week, but you can still join.

Statistical Learning

The Statistical Learning course is an introductory-level course in supervised learning with a focus on regression and classification methods. The course uses R, and the text book, “An Introduction to Statistical Learning with Applications in R”, is currently available as a free download.

Mining Massive Datasets

The Mining Massive Datasets course teaches algorithms for extracting models and other information from very large amounts of data, with an emphasis on techniques that are efficient and scale well. The course textbook, “Mining of Massive Datasets”, is currently available as a free download.

Machine Learning

The Machine Learning course provides a broad introduction to machine learning, data-mining, and statistical pattern recognition.

Filed Under: Big data Tagged With: training

Test your knowledge of R

February 4, 2015 By editor

Think your R knowledge is good? Well, put your money where your mouth is and take the Base R Assessment. You’ll be presented with 20 random questions from a pool of 74 (currently). Once you have completed the test, you’ll receive a ranking showing where you sit compared to all the others who have taken the test.

If you score less than 10, then we’re here to help…

Filed Under: Data analysis Tagged With: R

  • « Previous Page
  • 1
  • …
  • 44
  • 45
  • 46
  • 47
  • 48
  • …
  • 59
  • Next Page »

Copyright © 2026 · Decision Mechanics Limited · info@decisionmechanics.com