Sunday, March 13, 2016

[DataScience] Data Scientist 炼成记录-欢迎参与

数据科学:简单说就是,不要靠拍脑袋下结论,要以数据为根据,让事实说话。
.鏈枃鍘熷垱鑷�1point3acres璁哄潧
能力范畴3个词:统计编程表述


A PhD Data Scientist: Jack of All trades, master of one.



展开说:统计(能探索数据,建模,设计实验),
编程(能取数据,洗数据,至少能Prototype自己的data solution,懂基本大数据工作原理(MapReduce)),
表述(化繁为简,口头Present,书面写报告和论文,作图(静态和web))

. From 1point 3acres bbs
简历上(+脑子里)如果有这些:你找工作基本没有问题:. 涓€浜�-涓夊垎-鍦帮紝鐙鍙戝竷
    Ttest, Regression, ANOVA, Logistic Regression, DOE, Machine Learning, Data Mining, MapReduce, SQL, R/Matlab, Python, Java


. 鐗涗汉浜戦泦,涓€浜╀笁鍒嗗湴

=========================================
本文主要针对IT类行业做数据科学 It does not define a data engineer. Rather, it's a close call to a "full-stack data scientist". Master this list and you will not only be able to work for established firms, but startups too. 
其他偏重传统行业应用的,应该对表述要求稍高,对其他要求稍低。. from: 1point3acres.com/bbs 
面试之前请务必花1周时间学习对方行业的基本内容,wikipedia即可,起码做到熟悉对方行业常用关键字。
如果目的就是有份还可以的工作,请照单子静下心学习。
如果你希望做的很好,三个方面请突出至少一个方面。
要学过来,需要很多时间,如果希望不太费力就做data scientist, OK, dream on!

请不要mark一份学习清单就.Equals(学习任务已经完成了)一样,一起来学起来吧~~~~~~
【墙裂建议贴出你的学习计划,大家一起监督讨论,几位版主有空也会来给建议,坚持下来的有积分奖励】=========================================

如果有不清楚的请多google.

=========================================. 1point 3acres 璁哄潧
差不多一年前看市面工作还是很混杂的样子,今天又翻了翻,估计年底账目清算,很多公司很多新职位出来了,职位要求解析在此
感觉现在data scientist/researcher之类职位针对性更强,能更清楚看出来到底对方需要的是什么样的人:是啥都会一点的,还是会点统计的码农,还是Machine learning,还是优化、logistics 供应链,还是会点编程的统计师。
(data business person 一般不叫data scientist) 主要用SQL产生报表的BI analyst 也不在此列。

学习列表一来是准备面试用,二来本来平时就是要用的。我自己学完的mark as green
=========================================. 鐣欏鐢宠璁哄潧-涓€浜╀笁鍒嗗湴
打算把我自己学的一些东西总结在这里欢迎补充。不定期汇总到首楼。
如果你想收藏本帖请点首楼下方的“收藏” -》 确定 -》 然后文章会出现在 “快捷导航”-》收藏里面. 鍥磋鎴戜滑@1point 3 acres
如果没有啥具体内容要补充的,请不必回帖了。想加分的可以加分,不加也无所谓。

请别问我某校的Data Science项目如何,你三围如何能否上某校。I have no idea. .1point3acres缃�

=========================================
基本上是must have:

统计Statistics 统计和机器学习
  hypothesis testing, point/interval estimation
  pvalue, power, (type 1/2 error).鐣欏璁哄潧-涓€浜�-涓夊垎鍦�
  clt, delta method, derive coef and var(coef) etc 
  t-test: assumptions, remedy. 适用问题范围basics listed above 请看这个课 http://onlinestatbook.com/2/index.html
  glm (lm, logistic regression, anova etc):asssumptions, model selection and validation, diagnostics, remedy 适用问题范围

  times series         Forecast with R
         Time Series Analysis and Its Applications: With R Examples (Springer Texts in Statistics) 
         and its Upitt course

  bayesian
         Bayesian for hackers (python)
         Coursera Graphical Model (VERY nicely explained)
         Bayesian reasoning and machine learning book (quite difficult to read).1point3acres缃�
         入门:A first course in Bayes 一下就看完了,很不错

  longitudinal, mixed model. 鐣欏鐢宠璁哄潧-涓€浜╀笁鍒嗗湴
  doe:all kinds of design, response surface.1point3acres缃�
  (?)survival. 涓€浜�-涓夊垎-鍦帮紝鐙鍙戝竷

Machine Learning        Coursera Andrew Ng
    stanford Statistical Learning (Tibshrani & Hastie)
        -- 本书还出了一个本科版,着重动手实践,大量R, very easy to read. recommend starting from here. 
    Caltech那个learning from Data我没能跟下来Please, make sure you know your logistic regression inside and out!.1point3acres缃�

Learn recommender system
Learn some NLP 
Make sure you KNOW how things work, not just how to call a certain package in a certain language!!!
.鐣欏璁哄潧-涓€浜�-涓夊垎鍦�
Experimental Design / Causal InferenceThis is somewhat a niche area. But as a DS, you will most likely deal with some AB tests, if you are with a reputable internet company. It is not just using some tool to compute power for a chisquare test or t-test. Be sure you know the difference between observational study and designed experiment. Be sure you know when to use which.

统计软件Statistical Computing: R/Matlab/Python. SAS(?)
    R and Matlab 基本被业界认为是等同的。不过Matlab is not free, Octave is free 但是不是那么好用。请考虑自学R。反正你会Matlab 的话pick up R 也就分分钟的事情。
    如果其他语言一个都不会,只会SAS Base/Stat,并且你也不想学其他的,那也许数据科学不适合你。如果你非要用SAS不可,请你至少写过macro。SAS的确在大数据的建模里面非常有用,但是跟其他行业差距较大,如果组里其他人都是R/Py/Java 你跟他们交流起来会异常困难。另外软件很贵,很多地方未必愿意买。
    注意,我说的是,会SAS是好事,但是不能仅仅只会SAS. 
    Python: Data Analysis with Python (book), pandas
    R: data.table, or plyr, lubridate, reshape2, build a R package, there are now lots of such courses on both udacity and coursera. Start from any. 
        know how to get data from any source (DB, web, xml, plain text, etc)
EDA (exploratory) - Descriptive stats udacity
Inference - udacity
Plot/explain.鐣欏璁哄潧-涓€浜�-涓夊垎鍦�
read code from your favorite packages

-----------------------------------------------------
编程 : A compiled language, and a scripting language
Python 
    我比较偏好Udacity一遍教一遍做quiz 的方式,光做题不讲(codecademy)我自己好像学不清楚
    Udacity CS101. visit 1point3acres.com for more.
    Udacity CS 215 (Algorithm, 比Coursera Princeton and Stanford要简单,快速过一遍不错).鏈枃鍘熷垱鑷�1point3acres璁哄潧
    Udacity (Peter Norvig) CS212 Design of a Computer Program 非常好,强烈推荐

Java 数据结构和算法
1. Udacity java (这门课我花了40小时学完)适合连什么是函数什么是赋值都不知道的人。
2. Data structure 数据结构建议必学       python: Problem Solving with Algorithms and Data Structures)
    Java:  Berkeley 61B http://www.cs.berkeley.edu/~jrs/61b/
        教材是Head First Java & Data Structures and Algorithms in Java,
       my progress bar: week 5, lab1, hw1.
3. Algorithm:                  Udacity Algo in Python 比较laid back,如果不太希望费劲,可以上这个课,不过还是严肃点好。。。
       Java Coursera Algo I&II (Princeton),如果对这个话题有兴趣,. 1point 3acres 璁哄潧
                  不限语言 Stanford Algo I&II也很好,两者不可相互代替。

很少会有人学的第一门语言是C#,所以C#还真没有什么特别入门的书,不推荐。如果没从前没学C, java, C++直接看C#的书简直无法理解
C++比较难,对data scientist 来说应用也没有java广。当然如果你是大牛,plz当我没说。. more info on 1point3acres.com

Design pattern:地里同学推荐的:
http://courses.caveofprogramming ... ns-and-architecture
https://www.youtube.com/playlist?list=PLF206E906175C7E07

根据我组里面试别人,和我在其他地方面试,量化一下:数科的编程到底需要什么水平?
我假定你有了上述其他的全部功底,除非职位特别强调是统计师,或者叫Data scientist, statistics/analytics,并且职位说明里面对代码完全一带而过,你都可以假设,是需要一些代码能力的 。
具体水平是:
IT公司数科:Leetcode Medium要可做。所以,刷题吧。
传统公司:不知道
如果你是码农出身,或者做更偏向data engineer的,要求会更高

涉及知识点包括并且不限于:
     浮点溢出
     边界情况考虑. 1point3acres.com/bbs
     改进MapReduce算法(beyond brute force)
     如果涉及大数据,对时间复杂度要求会比较高
-- 其他我想起来了慢慢补

顺手学掉的小零碎:
Regex (a couple of hours) http://deerchao.net/tutorials/regex/regex.htm



SQL (a week) http://www.w3schools.com/sql/    Coursera: Intro to DB
大数据:. from: 1point3acres.com/bbs 

MapReduce: some knowledge    Udacity series:    http://blog.udacity.com/2013/11/sebastian-thrun-launching-our-data.html   
Coursera: intro to Data Science  
    Coursera: Big data and web intelligence. From 1point 3acres bbs    learning by doing --- yes! wrote my very first reducer for real life projects!    MongoDB (udacity) (NOSQL)
. Waral 鍗氬鏈夋洿澶氭枃绔�,
Spark/Scala - try this book: Advanced Analytics with Spark http://shop.oreilly.com/product/0636920035091.do (very doable and easy to follow, superb examples)
Scala推荐 Coursera: functional programming in scala - 超级好

Spark MOOC http://www.1point3acres.com/bbs/thread-135600-2-1.html
Book: Learning spark

If your want to be a DS for IT firms, then Maybe:
   jquery/ajax (start from codecademy very simple js and jquery intro, then find books) w3c school one is also really good.
-----------------------------------------------------
web services   get basic idea of how browsers work (udacity - Website Performance optimization)
   udacity web development (build a blog) (40 hours).鏈枃鍘熷垱鑷�1point3acres璁哄潧
-----------------------------------------------------
SE
   Software Development Life Cycles (udacity, mostly videos, as a quick intro only), amazingly, this one filled lots of holes in my knowledge base. Highly recommend. 鐣欏鐢宠璁哄潧-涓€浜╀笁鍒嗗湴
   Also a book is mentioned here, worth a quick flip through, unfortunately, no ebook that I found works. Martin Fowler, Kent Beck, John Brant, William Opdyke, Don Roberts-Refactoring_ Improving the Design of Existing Code.1point3acres缃�

-- this is helpful not only for working in IT, but helps overall coding style/efficiency as well. Wished I'd known earlier.
-----------------------------------------------------
Linux
   Many servers are in linux. at least familiarize yourself with the command line stuff. There's a not so good course on Edx. 
Basic shell script or similar 鏉ユ簮涓€浜�.涓夊垎鍦拌鍧�. 
jq, sed, awk.鐣欏璁哄潧-涓€浜�-涓夊垎鍦�
-----------------------------------------------------
综合/分析/表述/软技能
    软技能难以表述,
技巧不是最重要,想清楚再开口才是关键。突然发现我导师的lab页面竟然是用这些问题开头,深感心有戚戚。


化繁为简,高屋建瓴的表达能力:hide complex formula/engineering details,尽量传达big picture
    个人经验是,习得这些能力最好的办法是:去讲,不要自顾自的讲话,请随时关注听众是否听懂,鼓励对方马上提问,回答问题要选取符合对方背景的关键字,而不是“自己熟悉”的关键字。不要用缩写,小范围术语。多讲清楚intuition,少堆积公式。
    1. 教一门自己专业的入门课,e.g 统计学生,去给其他专业的人讲入门统计,例子:请给完全不懂统计的人讲,什么是pvalue, power, false positive, randomization, inference etc. 
    2. Consulting - 有些学校会有这种session,别觉得浪费时间,去把别人讲懂,去看看别人用你的专业技术做什么问题,他们的思路跟你哪里不同,你如何理解他们,如何让他们理解你。
    3. 做presentation - 不要像专业学术会议上那样去讲,要向给别人上101课那样讲。讲的目的,不是展示你的专业多么复杂深奥,不是为了impress others with your techinal prowess,而是让对方懂,最终听取你的建议。
    Data Journalism (course, starting early 2014) --- it was not as good as I expected. I do not recommend it. . 1point3acres.com/bbs


作图,静态的最好能会ggplot (a few hours), 动态的d3,如果你会javascript, also great!, 推荐读
     Nathan Yau: books visualize this & Data points, and his flowing data blog
     for d3: Interactive Data Visualization for the Web . free online tutorial by author: http://alignedleft.com/tutorials/d3/about 真的没那么难
    作图是否好看并不是关键所在,选用合适的图标来帮助解释道理才比较重要
html (a few hours, w3c)
css (a few hours, w3c), or codecademy, or the d3 book mentioned above. 1point3acres.com/bbs
javascript (codecademy as a start, a book to follow later)

Rcharts/highcharts
Udacity现在也有一门新开的vis课了
. 1point3acres.com/bbs
Prototype your data products: 
    mean stack. https://thinkster.io/angulartutorial/mean-stack-tutorial/
    起码把AngularJS学了,这个不光做数科有用。
    R open CPU. R Shiny (limited usage with free version).     If you are not into Angular, try the flask+React stack, 上手的确很快
  (关于flask, udacity有课,react自学即可,可以参考udacity 关于components的课)

虽然我们不是要做前段开发,但是看起来也得至少有个半吊子前段,请学习这MM的经验,超赞 http://www.1point3acres.com/bbs/thread-104335-1-1.html. 鐗涗汉浜戦泦,涓€浜╀笁鍒嗗湴
Design:  (optional but nice to know) 如果没有兴趣请至少看(组合起来好看的颜色)  如果你有兴趣让图好看,请花一个周末翻看这几本:
    1. Before and After
    2. Nondesigner's design book.鐣欏璁哄潧-涓€浜�-涓夊垎鍦�
    3. Don't make me think
    4. The Wall Street Journal Guide to Information Graphics

Research/publication:
    sharelatex (invite enough users to get free versioning) /writelatex.com
    Go to conferences, see what people are working on. Read their papers. 
    如果你想找某些类型的工作,上linkedin找到组员,泛读他们的paper

Domain Knowledge: google/wikipedia is your friend-google 1point3acres

=========================================
整体思路:
    Doing Data science (book) 鏉ユ簮涓€浜�.涓夊垎鍦拌鍧�. 
    Data Science in Business
=========================================
other 一些我感觉不太费时间但是会有用的小东西
   excel, power pivot etc. 鍥磋鎴戜滑@1point 3 acres
   科普类的书:(都很简单易读)

大数据到底是啥???http://www.amazon.com/Big-Data-Revolution-Transform-Think-ebook/dp/B009N08NKW/ref=sr_1_1?ie=UTF8&qid=1384931538&sr=8-1&keywords=big+data
和很近似的一本 http://www.amazon.com/Automate-This-Algorithms-Markets-World-ebook/dp/B0064W5UAS/ref=sr_1_8?ie=UTF8&qid=1384931546&sr=8-8&keywords=algorithms
随便翻翻就好了
然后当然还有Nate Silver http://www.amazon.com/The-Signal-Noise-Predictions-Fail-but-ebook/dp/B007V65R54/ref=pd_sim_kstore_1
=========================================
Case study:  Twitter data analytics http://tweettracker.fulton.asu.edu/tda/
=========================================
有人推荐的 MS  data science 学习curriculum  http://datasciencemasters.org/
=========================================. 1point 3acres 璁哄潧
大家给我推荐的帮助整理思路,用正确的方式做事的工具:It's more important than you think!!
http://software-carpentry.org/lessons.html
coursera reproducible research,学转knitr,不要copy paste anything. 1point 3acres 璁哄潧

Udacity Git Course (最好,没有之一)
============================
最后,没有什么比亲自干活和得到feedback更有用。
数据科学是一种 apprenticeship model,找合适的人带着做事,成长会很快。

Friday, March 11, 2016

Master's in Data Science

http://www.mastersindatascience.org/

Data Science in Energy

Opportunities in Energy Data Science

The Promise of Big Data

The energy industry is awash in data. Information streams in from a dizzying array of sources – exploration, production, transportation and distribution – and businesses are struggling to organize it.
What’s more, these companies are juggling expensive technologies while fending off smaller error margins, cutthroat competition and tighter government regulations. Easy energy is a long-abandoned dream. Wherever they fall on the creation-to-consumption line, companies need all the assistance they can get.
Analyzed correctly, big data has the potential to help the industry:
  • Discover new energy sources
  • Save money on drilling and exploration
  • Increase efficiency and productivity
  • Predict and stop accidents before they happen
  • Avoid power outages
  • Gauge consumption patterns
  • Match supply to demand
  • Plan for better maintenance and repairs
And, of course, improve profit margins and long-term viability.

Exploration and Discovery

So you’ve got 2D, 3D, 4D seismic monitoring data points. Now you gotta make something of ’em.
First, of course, you need to determine where to look for new oil or gas fields. You may even notice potentially productive areas via seismic trace signatures right in your current fields.
Once you’ve made your discovery, you’ll need to assess the likelihood that it will be profitable. Multiple parallel processing platforms are now used to process the host of data variables that can affect the viability of drilling operations:
  • Soil quality
  • Geologic anomalies
  • Production costs
  • Weather-related factors
  • Transport considerations
  • And more
This analysis can help you estimate how much oil or gas is left to be extracted. It all comes down to data: a well’s historical production along with local drilling, weather and environmental data (e.g., ocean currents for offshore rigs). The result is a much clearer picture of what you have to work with.

Digital Oilfields

Imagine an oilfield where every single piece of equipment was relaying a constant stream of data back to headquarters. Smart rigs would inform operators so they can maintain production flows. Sensors would alert workers about wells in need of repairs. Compressors would warn staff when they’re in danger of overloading.
Think it’s a bit sci-fi? Wrong. It’s already happening. Chevron calls it the “i-field,” BP the “Field of the Future,” and Royal Dutch Shell a “Smart Field.” Many simply call it the digital oilfield.
By combining sensor information (e.g., pressure, temperature, volume, shock and vibration data) with real-time data analytics and high-speed international communications, oil companies can every step of the production process – from initial extraction to daily maintenance.
The aim is to squeeze every last drop of black gold from their investment. Citing industry estimates, Chevron has suggested that it could generate 8% higher production rates and 6% higher recovery rates from a “fully optimized” digital oilfield.
That means a lot of money for a multi-billion dollar company, and it all comes down to data. And somebody has to work with that data.

Accident Prevention

As anyone who has lived through Deepwater Horizon, Exxon Valdez or Fukushima will tell you, failures can cost lives. The hope is that energy data science can help to stop disasters before they happen.
Sensors help companies monitor the life cycle of each component and catch problems early:
  • Are pressure and temperature surging? Enact safety measures.
  • Is drilling about to shatter a fragile environmental barrier? Shut it down.
  • Will a key piece of equipment need new parts soon? Have them ready and waiting.
It also goes beyond sensors. Data can be harvested from weather forecasts, geologic surveys, maintenance reports, video feeds – you name it – to detect unusual patterns, identify red-light situations and create a clearer picture of risk.
  • Are keywords like “leakage” or “vibration” clustering in a single area? Zero in on the problem.
  • Have algorithms detected a security breach overseas? Locate the culprit.
In an industry where millions can be affected by a single mistake, these steps aren’t optional. They’re mandatory.

Clean Energy

Helping the environment can also help the bottom line, and certainly helps job prospects for data scientists. Renewable energy sources like wind, solar and tides are hot and getting hotter.
Take wind. Like oil rigs, wind turbines are made up of hundreds of moving parts. Sensors on these parts generate data on everything from wind speed to pitch and yawn degrees. They inform monitors if the turbine is working at peak performance, whether machinery needs maintenance, and if failure is imminent.
Add condition monitoring data (e.g., vibration monitoring, acoustic emissions), work order data (e.g., repair costs) and historical data (25 years and counting) and you have a lot of places to find efficiencies and cut costs.
Even utilities have something to smile about. Although wind and solar are notoriously fluctuating sources of energy, smart grids are helping companies manage the uncertainties. As tidal and solar improve in output and storage, so too will the big data technologies used to monitor and administer to their working parts.

Data Risks

The Challenges Ahead

So what’s stopping the U.S. from having the most efficient, data-driven energy industry in the world? A few things.

1. Disorganization

I’ll step aside and let Jamal Khawaja sum up the problem:
“Many organizations admit that they are not making the most of the information assets they have residing in structured repositories, and hardly any are exploiting the data held outside of structured systems to any significant degree. So let me frame the problem: it’s not Big Data that is confounding CIOs; it’s deriving value from that data.”
Energy companies know they have huge volumes of data going to waste. They just don’t know how to handle it.
It’s not surprising. The volume, velocity, veracity and variety (the 4Vs) of big data – especially in an industry with so many moving parts – can stagger any data scientist. Finding ways to transform this information into actionable insights is not going to be easy.

2. Variations in Data

Energy has a related problem. Its data sources are all over the map. Here’s Khawaja again:
“In the oil and gas industry, only a subset of data exists in a format that can be easily ingested by a relational database. The majority of data collected from wells, operations, and other instrumented locales exist in tagged, flat-file format or organized according to XML-based standards designed to conform to the aggregating software.
‘Because there is such detailed characterization of wells, equipment, facilities and other entities from an instrumentation perspective, there is no easy way to convert this complex, unstructured data into a relational database format.”
Data science is rapidly finding ways to overcome these problems, but for the moment, corralling diverse sources remains a challenge.

3. Lack of Data Expertise

This may be the biggest problem. And a most interesting one, too, for career-seekers in the next 20 years. You see, the energy industry is going to need a lot of smart analysts to help uncover answers lurking in their data.
“By 2018, the United States alone could face a shortage of 140,000 to 190,000 people with deep analytical skills as well as 1.5 million managers and analysts with the know-how to use the analysis of big data to make effective decisions.”
That’s a big shortage. It means good work for a lot of math-lovers who also understand business and who know how to communicate. What’s more, it’s a position that will likely come with a large paycheck and a giant burden of responsibility. It will confer prestige and status on its holders, both in their organizations and in society.
After all, we’re not just talking about boosting a retail store’s quarterly profits. This is about keeping the lifeblood of the future flowing smoothly.

History of Data Analysis and Energy

“Energy forecasting is easy. It’s getting it right that’s difficult.” – Graham Stein

On August 27, 1859, after a series of setbacks and frustrations, Edwin Drake and his driller, Billy Smith, halted work for the day. Having punched through gravel and over thirty feet of bedrock, their drill bit was now over 69 feet below the surface.
Drake had been hired by Seneca Oil of Connecticut to investigate intriguing deposits near Titusville, Pennsylvania. He, in turn, employed Smith, an expert in drilling for salt.
It could have come to naught. But on the morning of August 28, Smith saw something he’d never seen there before. Something that would spark a boom in drilling, entrepreneurship and data science over the next century: a bubbling sludge of crude oil.
What was so special about that puddle of gurgling goop? Blame it on improving technology and economics.

The Wild Years of Exploration and Discovery

During the late 19th century, oil was in high demand. Whale oil, once the favored fuel for lamps, had been usurped by cheaper kerosene. The country demanded light.
But in those days, striking oil was a hit-and-miss prospect. Methods were haphazard. Many petroleum prospectors sunk their wells in places near known oil and gas seeps and hoped for the best.
Even so, there were plenty of discoveries to go around. When the Civil War intervened in the flow of oil from the east, companies simply migrated west to California and south to Texas, Oklahoma, Louisiana and Arkansas.
As the new century dawned, demand remained insatiable. The oil industry turned its attention from kerosene to gasoline (once considered a useless byproduct) for automobiles and airplanes. There was a lot of money, a lot of guesswork and a lot of waste.

A Seismic Shift in Data Collection

Leading up to and during the 1920s, oil companies began to realize that science could solve some of their problems. Geologists like Wallace Pratt and J. Clarence Karcher were employed to conduct studies on potential oil fields and investigate new tools and techniques.
One of these tools was the seismograph. Originally developed to monitor earthquakes, the seismograph had another handy purpose. By generating small explosions, typically with dynamite, geologists could detect how seismic waves were behaving under the surface of the earth.
From the data collected, scientists could then create a detailed map of the subsurface, including the shape and position of underground rock layers. Combined with data from surface geologic tests, this “seismic reflection profile” would give companies a much better sense of where to drill.
The era of big data in energy had begun.

Farewell Slide Rules, Hello Computers

World War II brought with it new uses for petroleum and natural gas products (e.g., TNT and artificial rubber) as well as the initial development of supercomputers. By this time, oil companies were busily hoarding data from an increasing variety of investigative engineering techniques, including magnetometers and well logging.
Still, it wasn’t until the 1960s and 1970s that data analysis in the energy industry really began to heat up:
  • Data processing algorithms for velocity analysis, refraction, residual statics correction and stacking, deconvolution and migration were developed.
  • Mainframe computers running complex drilling computations replaced slide rules and calculators.
In 1963, Humble Oil Company developed new 3D seismic technology. According to Exxon, this data-heavy technology – coupled with the use of massively parallel computers in seismic imaging – helped to sharply reduce finding costs from the 1980s onward.
At the same time, the renewable energy sector was beginning to find its feet. Recognizing future needs, the government and start-up companies began to experiment with wind and solar. The first large commercial electricity-generating wind turbines appeared in the 1970s. And naturally, all that research and development generated even more data.

New Technologies, New Booms

Meanwhile, computers were shrinking, but their power was increasing exponentially. Developments in hardware and software, and improved graphics – not to mention the advent of the Internet – enabled engineers to contemplate boldly going where no drill had gone before.
This was good news for petroleum companies, operating in a worldwide landscape where demand was high but supply appeared to be dwindling.
By the mid-1990s, scientists working as part of a Statoil-Schlumberger joint project were able to use 4D seismic monitoring – analysis of 3D seismic data captured at different times in the same area of an oil field – to differentiate between drained and undrained areas and identify the remaining pockets of oil and gas.
Other searchers applied technical advancements in horizontal drilling and hydraulic fracturing to extract natural gas. In 1997, after a series of failed experiments, Mitchell Energy completed the first economically viable fracture of Texas’s Barnett Shale using slick-water fracturing. This touched off a U.S. rush on natural gas, as well as an avalanche of data, that have yet to subside.

Wednesday, March 9, 2016

Data Science Use Cases

Background

For each type of analysis think about:
  • What problem does it solve, and for whom?
  • How is it being solved today?
  • How can it beneficially affect business?
  • What are the data inputs and where do they come from?
  • What are the outputs and how are they consumed- (online algorithm, a static report, etc)
  • Is this a revenue leakage ("saves us money") or a revenue growth ("makes us money") problem?

Use Cases By Function

Marketing

  • Predicting Lifetime Value (LTV)
    • what for: if you can predict the characteristics of high LTV customers, this supports customer segmentation, identifies upsell opportunties and supports other marketing initiatives
    • usage: can be both an online algorithm and a static report showing the characteristics of high LTV customers
  • Wallet share estimation
    • working out the proportion of a customer's spend in a category accrues to a company allows that company to identify upsell and cross-sell opportunities
    • usage: can be both an online algorithm and a static report showing the characteristics of low wallet share customers
  • Churn
    • working out the characteristics of churners allows a company to product adjustments and an online algorithm allows them to reach out to churners
    • usage: can be both an online algorithm and a statistic report showing the characteristics of likely churners
  • Customer segmentation
    • If you can understand qualitatively different customer groups, then we can give them different treatments (perhaps even by different groups in the company). Answers questions like: what makes people buy, stop buying etc
    • usage: static report
  • Product mix
    • What mix of products offers the lowest churn? eg. Giving a combined policy discount for home + auto = low churn
    • usage: online algorithm and static report
  • Cross selling/Recommendation algorithms/
    • Given a customer's past browsing history, purchase history and other characteristics, what are they likely to want to purchase in the future?
    • usage: online algorithm
  • Up selling
    • Given a customer's characteristics, what is the likelihood that they'll upgrade in the future?
    • usage: online algorithm and static report
  • Channel optimization
    • what is the optimal way to reach a customer with certain characteristics?
    • usage: online algorithm and static report
  • Discount targeting
    • What is the probability of inducing the desired behavior with a discount
    • usage: online algorithm and static report
  • Reactivation likelihood
    • What is the reactivation likelihood for a given customer
    • usage: online algorithm and static report
  • Adwords optimization and ad buying
    • calculating the right price for different keywords/ad slots

Sales

  • Lead prioritization
    • What is a given lead's likelihood of closing
    • revenue impact: supports growth
    • usage: online algorithm and static report
  • Demand forecasting

Logistics

  • Demand forecasting
    • How many of what thing do you need and where will we need them? (Enables lean inventory and prevents out of stock situations.)
    • revenue impact: supports growth and militates against revenue leakage
    • usage: online algorithm and static report

Risk

  • Credit risk
  • Treasury or currency risk
    • How much capital do we need on hand to meet these requirements?
  • Fraud detection
    • predicting whether or not a transaction should be blocked because it involves some kind of fraud (eg credit card fraud)
  • Accounts Payable Recovery
    • Predicting the probably a liability can be recovered given the characteristics of the borrower and the loan
  • Anti-money laundering
    • Using machine learning and fuzzy matching to detect transactions that contradict AML legislation (such as the OFAC list)

Customer support

  • Call centers
    • Call routing (ie determining wait times) based on caller id history, time of day, call volumes, products owned, churn risk, LTV, etc.
  • Call center message optimization
    • Putting the right data on the operator's screen
  • Call center volume forecasting
    • predicting call volume for the purposes of staff rostering

Human Resources

  • Resume screening
    • scores resumes based on the outcomes of past job interviews and hires
  • Employee churn
    • predicts which employees are most likely to leave
  • Training recommendation
    • recommends specific training based of performance review data
  • Talent management
    • looking at objective measures of employee success

Use Cases By Vertical

Healthcare

  • Claims review prioritization
    • payers picking which claims should be reviewed by manual auditors
  • Medicare/medicaid fraud
    • Tackled at the claims processors, EDS is the biggest & uses proprietary tech
  • Medical resources allocation
    • Hospital operations management
    • Optimize/predict operating theatre & bed occupancy based on initial patient visits
  • Alerting and diagnostics from real-time patient data
    • Embedded devices (productized algos)
    • Exogenous data from devices to create diagnostic reports for doctors
  • Prescription compliance
    • Predicting who won't comply with their prescriptions
  • Physician attrition
    • Hospitals want to retain Drs who have admitting privileges in multiple hospitals
  • Survival analysis
    • Analyse survival statistics for different patient attributes (age, blood type, gender, etc) and treatments
  • Medication (dosage) effectiveness
    • Analyse effects of admitting different types and dosage of medication for a disease
  • Readmission risk
    • Predict risk of re-admittance based on patient attributes, medical history, diagnose & treatment

Consumer Financial

  • Credit card fraud
    • Banks need to prevent, and vendors need to prevent

Retail (FMCG - Fast-moving consumer goods)

  • Pricing
    • Optimize per time period, per item, per store
    • Was dominated by Retek, but got purchased by Oracle in 2005. Now Oracle Retail.
    • JDA is also a player (supply chain software)
  • Location of new stores
    • Pioneerd by Tesco
    • Dominated by Buxton
    • Site Selection in the Restaurant Industry is Widely Performed via Pitney Bowes AnySite
  • Product layout in stores
    • This is called "plan-o-gramming"
  • Merchandizing
    • when to start stocking & discontinuing product lines
  • Inventory Management (how many units)
    • In particular, perishable goods
  • Shrinkage analytics
    • Theft analytics/prevention (http://www.internetretailer.com/2004/12/17/retailers-cutting-inventory-shrink-with-spss-predictive-analytic)
  • Warranty Analytics
    • Rates of failure for different components
      • And what are the drivers or parts?
    • What types of customers buying what types of products are likely to actually redeem a warranty?
  • Market Basket Analysis
  • Cannibalization Analysis
  • Next Best Offer Analysis
  • In store traffic patterns (fairly virgin territory)

Insurance

  • Claims prediction
    • Might have telemetry data
  • Claims handling (accept/deny/audit), managing repairer network (auto body, doctors)
  • Price sensitivity
  • Investments
  • Agent & branch performance
  • DM, product mix

Construction

  • Contractor performance
    • Identifying contractors who are regularly involved in poor performing products
  • Design issue prediction
    • Predicting that a construction project is likely to have issues as early as possible

Life Sciences

  • Identifying biomarkers for boxed warnings on marketed products
  • Drug/chemical discovery & analysis
  • Crunching study results
  • Identifying negative responses (monitor social networks for early problems with drugs)
  • Diagnostic test development
    • Hardware devices
    • Software
  • Diagnostic targeting (CRM)
  • Predicting drug demand in different geographies for different products
  • Predicting prescription adherence with different approaches to reminding patients
  • Putative safety signals
  • Social media marketing on competitors, patient perceptions, KOL feedback
  • Image analysis or GCMS analysis in a high throughput manner
  • Analysis of clinical outcomes to adapt clinical trial design
  • COGS optimization
  • Leveraging molecule database with metabolic stability data to elucidate new stable structures

Hospitality/Service

  • Inventory management/dynamic pricing
  • Promos/upgrades/offers
  • Table management & reservations
  • Workforce management (also applies to lots of verticals)

Electrical grid distribution

  • Keep AC frequency as constant as possible
  • Seems like a very "online" algorithm

Manufacturing

  • Sensor data to look at failures
  • Quality management
    • Identifying out-of-bounds manufacturing
      • Visual inspection/computer vision
    • Optimal run speeds
  • Demand forecasting/inventory management
  • Warranty/pricing

Travel

  • Aircraft scheduling
  • Seat mgmt, gate mgmt
  • Air crew scheduling
  • Dynamic pricing
  • Customer complain resolution (give points in exchange)
  • Call center stuff
  • Maintenance optimization
  • Tourism forecasting

Agriculture

  • Yield management (taking sensor data on soil quality - common in newer John Deere et al truck models and determining what seed varieties, seed spacing to use etc

Mall Operators

  • Predicting tenants capacity to pay based on their sales figures, their industry
  • Predicting the best tenant for an open vacancy to maximise over all sales at a mall

Education

  • Automated essay scoring

Utilities

  • Optimise Distribution Network Cost Effectiveness (balance Capital 7 Operating Expenditure)
  • Predict Commodity Requirements

Other

  • Sentiment analysis
  • Loyalty programs
  • Sensor data
    • Alerting
    • What's going to fail?
  • De duplication
  • Procurement

Use Cases That Need Fleshing Out

Procurement

  • Negotiation & vendor selection
    • Are we buying from the best producer

Marketing

  • Direct Marketing
    • Response rates
    • Segmentations for mailings
    • Reactivation likelihood
    • RFM
    • Discount targeting
    • FinServ
    • Phone marketing
      • Generally as a follow-up to a DM or a churn predictor
    • Email Marketing
  • Offline
    • Call to action w/ unique promotion
    • Why are people responding- How do I adjust my buy (where, when, how)?
    • "I'm sure we are wasting half our money here, but the problem is we don't know which ad"
  • Media Mix Optimization
    • Kantar Group and Nielson are dominant
    • Hard part of this is getting to the data (good samples & response vars)

Healthcare

  • CRM & utilization optimization
  • Claims coding
  • Forumlary determination and pricing
  • How do I get you to use my card for auto-pay? Paypal? etc. Unsolved.
  • Finance
    • Risk analysis
    • Automating Excel stuff/summary reports