Tag Archives: data science

A Data Science Career with Kirk Borne, Free Webinar

Once again, The Data Incubator, is hosting another Data Science in 30 minutes webinar. This one features the career of Kirk Borne.

Renowned data scientist, Kirk Borne will take viewers on a journey through his career in science and technology explaining how the industry-and himself have evolved over the last 4 decades. Starting with skipping lunches in high school to a systematic twitter obsession, Kirk will shed light on his road to success in the data science industry.

Kirk is universally considered one of the most (if not the most) influential voices in data science. If you are interested in a career in data science, this is a webinar you will not want to miss.

The webinar is 5:30 Eastern Time on August 29, 2017, and registrations are currently being accepted. It is free.

The 3 Stages of Data Science

Businesses everywhere are racing to extract meaningful insight from their data. Many organizations are spinning up data science teams and attacking problems (some more successful than others). However, one of the challenges is determining the current stage of data science within the organization. Next is determining the desired stage of data science.

Below are 3 stages of a truly mature data science organization.

1. Dashboards

The beginning stage of data science is dashboards. It is all about answering “How much?” and “What happened?” by looking at reports of historical data. If done well, it might even help an organization answer “Why”. Many organizations will refer to this phase as Business Intelligence.

The dashboard stage can be very expensive for an organization, in terms of people-hours and money. It usually involves investments in:

  1. Data Warehouse or some other storage environment, for storing the data in a single location for easy reporting
  2. ETL (Extract Transform Load) Tools for manipulating, combining, and moving data to the data warehouse
  3. Reporting Tools for displaying the results and allowing users to “explore” the data

Here are some common questions that can be answered via traditional dashboards:

  • How many customers live in each region?
  • What were the sales on Black Friday?
  • How many patients visited the hospital last month?

As you can see, there are large amounts of value that can be gained by this phase alone. It is critical for a business to clearly understand past performance. Unfortunately, this phase is where many businesses stop.

2. Machine Learning

The real “science” of data science does not begin until the second stage which is machine learning. It focuses on estimating quantities that cannot be directly observed. This could be what movies a customer will like, the price of a company’s stock tomorrow, or the causal effect of a particular advertising campaign. Machine Learning uses the data from the first phase and applies statistical or other methods to gain additional insights.

Think of machine learning as answering the following:

  • When a customer moves, will he/she spend money at a hardware store?
  • When a credit card purchase is made, what is the probability the charge was fraudulent?
  • What is the expected lifetime value of a new customer?
  • If a hurricane is coming, what will people buy? (pop tarts? it is true).

Notice the connection between an event and some outcome. The value of machine learning comes from estimating the causal outcome of potential events. This phase is filled with terms such as: machine learning, data mining, and statistical modeling. The machine learning stage is all about looking into the future!

3. Actions

Determining the actions to perform, is the third and final phase. It tries to capitalize on the results of machine learning in order to take appropriate actions. The following actions might be suitable for the events identified in the predictive section above.

  • When a customer moves, send a “welcome to the neighborhood” packet with coupons to nearby hardware stores.
  • Decline the fraudulent charge or deactivate the credit card.
  • If the new customer has very high expected lifetime value, provide some special treatment or offers to ensure the customer becomes a customer for life.
  • When a hurricane is approaching, place Pop tarts near the front of the store.

As you can see, good machine learning from the second phase can lead to clear actions.


Claiming success in Data Science is all about conquering all three stages. Each stage builds upon the previous stage. If you have put in the effort to complete the first stage, why not continue to the second and third stages?

Guidelines for Telling a Great Data Science Story

People love stories. People can connect with stories. People remember great stories. Make your data tell a story. If you can make stories come alive with data, people will pay attention.

There is no magic formula for a great story, data or otherwise. Here are some guidelines for telling a great data science story.

  • Clearly state the problem
  • Explain the data
  • Share the struggles of doing the analysis
  • Do not focus on the algorithms
  • Show how the analysis progressed, take your listeners on a journey
  • Finish with something remarkable

The late Hans Rosling could tell as good of a story with data as anyone. Do a quick internet search for his name, and you can easily find his Ted talks or other videos. He provides an excellent model for telling a story with data. It is worth your time to watch some of his videos.

The entire goal of telling a story with data is to get people engaged in the problem.

Leave a comment if you have others tips for telling an effective data science story.

Papers for Teaching Undergraduate Data Science

If you work at a university and are considering starting an undergraduate program in data science, then today’s post is for you.

If you know of any other papers, please leave a comment below.

Site For Undergraduate Data Science Programs

Karl Schmitt, Director of Data Sciences at Valparaiso University, has started a blog to share his experiences with building an undergraduate data science program. The blog is titled, From the Director’s Desk. Karl is regularly posting about textbooks, curriculum, visualizations and learning objectives from the perspective of an educator. Tons of great resources!

Valparaiso University is Turning Homework into Social Change

Recently, I had the honor of speaking with Dr. Karl Schmitt from Valparaiso University. He is the director of the Data Science undergraduate program at Valparaiso University. We had a very nice discussion, and I thought I would pass along my summary.

What are the Details of the Valparaiso Undergraduate Program?

The program is housed in the Mathematics department and it is designed to be fairly interdisciplinary. It consists of four parts.

  1. Math
  2. Statistics
  3. Computer Science
  4. A Separate Focus Area

The separate focus area can be from nearly any other department and is targeted at building some domain expertise. Although not required, a double major is encouraged.

One of the most unique and excited aspects of the program begins during the first year. Students take Introduction to Data Science, which has few prerequisites and serves as motivation for the remainder of the program. Valparaiso partners with non-profits and government agencies to provide the first year students with hands-on experience solving problems for social good. Examples include Meals on Wheels, mapping with the United States Geological Survey, and a child welfare non-profit. Then, the junior and senior students are involved with a capstone project that can be a continuation of the first year project, some other social good project, or students can serve in a consulting capacity to other departments on campus.

What skills Do You expect Valparaiso Data Science Graduates to Have?

There are a few basics skills that make sense for data science: coding, database skills, statistics, and general math. In addition, Valparaiso grads should also know how to talk, write, and create videos about mathematical concepts. Finally, ethics is an essential portion of the program. According to Dr. Karl Schmitt,

I want my students to graduate with ethics related to data science.

To enforce that statement: ethics case studies are required of all students, it is a key learning objective of the projects, and ethics is integrated into all the classes so students understand the importance. Students need to be able to do the hard data science, communicate the results and care about the consequences.

Why Choose Data Science as an Undergraduate?

It is a utility degree that is in strong demand in nearly every field. As companies continue to understand the usage of data, having data skills is going to get increasingly more crucial. Data Scientist are going to be (currently are) in demand for human resources, supply, sales, technology and many other awesome jobs.

Why Valparaiso for Data Science?

There are a number of reasons:

  • Good University Size – It is easy to double major and engage with things outside the major, plus disciplines are very connected which allows for collaboration.
  • Writing/Communication is Integrated Throughout – Many people can crunch numbers, but Valparaiso graduates can express discoveries. The students get that from the very beginning.
  • Projects – All students will have experience and examples of projects to demonstrate.
  • Finally, students have an opportunity to turn their homework into something that matters!

Thank you to Dr. Karl Schmitt for the interview and to Valparaiso University for Sponsoring Data Science 101.

Netflix Data Scientist on Machine Learning: Free Webinar

The Data Incubator, a data science fellowship program, is currently running a Data Science in 30 minutes webinar series. Next week features a free webinar with Dr. Becky Tucker of Netflix. Dr. Tucker is a Senior Data Scientist at Netflix where she specializes in predictive modeling for content demand (think what do people want to watch). The full abstract of the webinar is below. The webinar is free; all you need to do is register.

Predicting Content Demand with Machine Learning

Date/Time: March 9, 2017 @ 5:30 PM ET
Location: Online
Register: Click Here

Abstract: Netflix is well-known for its data-driven recommendations that seek to customize the user experience for every subscriber. But data science at Netflix extends far beyond that – from optimizing streaming and content caching to informing decisions about the TV shows and films available on the service. The talk will cover work done by Becky and the Content Data Science team at Netflix, which seeks to evaluate where Netflix should spend their next content dollar using machine learning and predictive models.

Update – Below is the Recorded Webinar

Building Data Science Skills as an Undergraduate

While there are a growing number of universities that offer undergraduate data science degrees, for one reason or another those programs may not be perfect for everyone interested in data science. So, what do you do if you attend a school that does not offer a data science degree? This is a question frequently asked of me, so I thought I would elaborate on my typical response.

You Cannot Know It All

First off, you will never know all there is to know about data science. The field is vast and contains many sub-fields. Thus, as an undergraduate, a good plan is to learn the fundamentals. Then expand your knowledge/expertise as your education and career continue. Data Science is evolving rapidly and it requires continual learning. Hopefully, this is one of the reasons you are interested in the field.

My Recommended Approach

A good plan is to major in computer science or statistics and minor in the other. If your school doesn’t have either of those major, then take as many of those classes as you can. Next, choose a domain specific area such as business, chemistry, psychology, etc.; and gear your elective classes toward that domain area. This approach will give you a solid base understanding of the statistical and computational underpinnings of data science. You should also be well-prepared to find a job or continue your studies in graduate school.

Also, somewhat related, taking an art class or two might not be a bad idea. Visualization is very important to data science. Understanding color palettes and usage of space on a canvas are concepts that will serve you well. Plus, many people strong in computer science and statistical algorithms are lacking in artistic skills.

Some Enhancements to Your Education

If your location allows, consider attending local meetups. Finally, get involved with whatever projects you can (Kaggle, internships, open source, …).

Do you have any advice for undergraduates looking to study data science? If so, please leave a comment.

Are you and undergraduate with questions? Please ask in the comments below.

Quora Answers by Monica Rogati

Monica Rogati, a legend in the data science space, recently provided some answers on Quora that are sheer internet gold.

Quora Answers by Monica Rogati

She answers questions involving:

  • What is a data science advisor?
  • Challenges of Building a data science team?
  • Characteristics of a good data scientist?
  • and more

They are filled with great advice.

Get me Sum ‘dat Big Data

I teach data science courses thoughout the US. I enjoying asking attendees why they are in class. I get many good answers, but occassionally I get some funny answers. Here is a story with one of the more humorous answers.

While chatting with an attendee before class, I asked why he chose to attend this class. Here was his answer.

Well, my boss attended a conference and heard a talk on Big Data. Then, he came back to the office and bought hadoop for some of our systems. Next he heard about this training and told me to attend. When preparing to leave, the boss said, “Get me sum ‘dat big data”.

After a slight chuckle from both of us, I mentioned we would talk more about that in class.

While this story is somewhat humorous, it is not all that uncommon. Companies want to start using data science, they often just do not know where to start. If you are looking for a starting point, check out this post, You Want Data Science, Now What?.

Do you have a funny “data science” or “big data” story? If so, please share in the comments.