Great quote on the use of hypothesis testing by @JPdeRuiter.
Reject a specific null, and then argue for an arbitrary alternative. It’s indeed pretty remarkable that so few people see how absurd this procedure is.
Insight. Applied.
By editor
Great quote on the use of hypothesis testing by @JPdeRuiter.
Reject a specific null, and then argue for an arbitrary alternative. It’s indeed pretty remarkable that so few people see how absurd this procedure is.
By editor
The Wall Street Journal ran an intriguing story last month about Netflix’s management overriding recommendations coming from the company’s algorithms.
Analysis showed that promotions for the comedy “Grace and Frankie” were more successful when they only featured one of the two stars of the show.
Apparently fearful of alienating one of their stars, Netflix’s management decided to include both in promotions—even though that would produce a sub-optimal response.
The subtitle of the article mentions
overriding the metrics
However, this isn’t how I see it.
Data science produces inputs to the decision-making process—not recommendations to be followed slavishly. Netflix’s management presumably considered all the information at their disposal and made a decision that they believed would maximise their long-term rewards.
This is as it should be…even at Netflix.
The formal analysis could have been extended to include information on the excluded star’s contract, longevity as an asset, propensity to be offended, etc.
Maybe game theory could have been applied…and some bright Netflix quant could have developed a “diva scale”. But, this would have complicated the analysis considerably and compromised its accuracy.
Looks like data and judgement might have been combined effectively in this case.
By editor
Google have created a tool to make it easier to discover datasets—Google Dataset Search.
One potential downside is that it requires dataset owners to provide metadata. While the Google brand means that this might get some traction, not all dataset owners are motivated to help the cause. Publication of datasets is now mandated by some funding bodies, but that doesn’t mean that the datasets have to be discoverable. Ideally we’ll see funding bodies now mandating the addition of metadata.
While we wait for Google Dataset Search to evolve, we can still rely on curated repositories and lists such as
By editor
Jupyter Notebooks are popular with data scientists. Microsoft even offers a free, hosted, “no-install” service for Python, R and F#.
However, there are some downsides to notebooks—mostly to do with software engineering best practices.
Joel Grus gave a provocative talk at JupyterCon 2018 entitled “I Don’t Like Notebooks”. Yihui Xie then followed up with a response to Grus’ talk.
Both authors make a good case and have interesting points. As ever, the truth is that notebooks are good in some situations and not so good in others.
Personally, I use both. Notebooks for smaller, exploratory, data science projects and IDEs (Visual Studio Code, PyCharm and RStudio) for everything else.
By editor
Python has topped the IEEE Spectrum list of top programming languages again this year—extending its lead in the process.
The sources used to compiled the list cover
contexts that include social chatter, open-source code production, and job postings.
Obviously that list of sources isn’t an accurate reflection of what developers are doing day-to-day in organisations. Any list of top programming languages that puts R (#7) above JavaScript (#8) clearly has some methodological challenges. My belief is that the list reflects the current buzz around data science.
However, interest in Python clearly remains high. As it does in R—#7 is impressive for a domain-specific language.