Data Science Interview Questions with Answers

by Thomas | Oct 29, 2018 | Data Science, Deep Learning, Learn Data Science, machine learning, Python, R Programming | 1 comment

Expertise Critical for Every Data Scientist

[embeddoc url=”https://dimensionless.in/wp-content/uploads/2018/10/Data-Science-topics.pdf” viewer=”google”]

The Best Way to Prepare for Interview Questions

Now suppose you read a question about a topic like overfitting. You can read the text and memorize the answer. Usually, articles with this heading (Interview Questions and Answers) are normally constructed that way, with plain text questions and answers. You could follow that route for interview preparation, but it is simply not the right thing to do. I can give you a list of important questions, with answers. Which is exactly what I will do in this article, later.

But you need to understand one thing clearly.

You cannot learn programming and data science from books alone.

You can learn the heading and the words. But the concept will truly be understood only in a practical manner; in a mini-project or in a worked-out example on the computer.

Data science is similar to programming in this regard.

Books are meant to just start your journey.

The real learning begins only when you implement it in code by yourself.

To take an example:

Question from the Interviewer:

“What is cross-validation and why is it important? How does it eliminate overfitting?”

A Good Answer:

“Cross-validation eliminates overfitting by exposing the model to the entire data set in a statistically uniform manner. Overfitting happens when the training set and test sets are not properly selected. If a model like LogisticRegression is trained until the error rate is very small, it may not be able to generalize to the pattern of data found in the test set. Hence the performance of the model would be excellent on the training set, but poor on the test set. This is because the model has overfitted itself to the training data. Thus, when presented with test data, error values increase because the generalization capacity of the model has been decreased and the model cannot discover the patterns of the test data.”

“K-fold Cross Validation prevents this by first dividing the total data into k sections and using one section as the test set and the remaining sections as the training set. We train k models, each time using a different fold as the test set and the remaining folds as the training set. Thus, we cover as many combinations of the training and test set as possible as input data. Finally, we take an average of the results of each model and return that as the output. So, overfitting is eliminated by using the entire data as input, one section (one of the k folds) being left out at a time to use as a test set. A common value for k is 10.”

Question:

“Can you show me how that works by coding it on a 10 by 10 array of integers? In Python?”

Worst Case Answer:

…

“Ummmmmmmm…..”

“Sorry sir, I just studied that in a textbook. I am not sure how I could work through that by code.”

(!!!)

You Can’t Study Without Implementation

Data science should be studied in the way programming is studied. By working at it on a computer and running all the models in your textbook, and finally, doing your own mini-project, on every topic that could be important. Can you learn to drive a car by reading about it in a book? You need practical experience! Otherwise, all your preparation is meaningless. That is the point I wanted to make.

Now, having established this, I assume from here on that you are a data scientist in training who has worked the fundamental details on a computer and is familiar with the basics. You just need the finishing touches on your interview preparation. If that is the case; here are your topics for mini-projects and experiments! And – interview questions with answers.

Interview Practice Resources

Python Practice

https://www.testdome.com/d/python-interview-questions/9

This is a site that allows you to sharpen your skills in Python for interviews. There are many more sites like these, all you need to do is Google ‘Python Interview Questions’.

R Practice

https://www.computerworld.com/article/2497143/business-intelligence/business-intelligence-beginner-s-guide-to-r-introduction.html

Many people know Python, but R is not as commonly known. The above tutorial spans 30 pages that you can work through with your R console to learn the basics. Alternatively, you could try Swirl (link given below), which is also highly recommended for beginners.

https://swirlstats.com/

Kaggle

Work through Kaggle competitions. No better way to establish yourself in the data science universe.

https://www.kaggle.com/competitions

Also, if you have basic data science skills, try your hand with the hands-on Kernels section. Cash prizes awarded every week!

https://www.kaggle.com/kernels

Oh, what are kernels? Kaggle Kernels are online Jupyter notebooks that allow you to run Python and R code interactively with your browser in the same application without any local processing. All computation is done on the Kaggle servers.

1. What is a normal distribution? And how is it significant in data science?

The normal distribution is a probability distribution, characterized by its mean and standard deviation or variance. The normal distribution with a mean of 0 and a variance of 1 looks like a bell, hence it is also referred to as the bell curve. The central limit theorem makes the normal distribution ubiquitous in data science. In its essence, the central limit theorem states that data values tend to be attracted to the normal distribution shape as the number of samples is increased without limit. This theorem is used in data science nearly everywhere, because it gives you an ‘expected’ value for an arbitrary dataset that has, say, n = one thousand samples. As n increases, if the data is normally distributed, the shape of the graph of that attribute will tend to look like the bell curve.

2. What do you mean by A/B testing?

An A/B test records the results of two random variables or hypotheses (depending upon the scenario) and compares the rate of success or accuracy for the variable being in the state of A or the state of B. This often tells us which feature should be used to build a machine learning model. It is also used to select which model to use in the first place. A/B testing is a general concept that can be applied to nearly every system.

3. What are eigenvalues and eigenvectors?

The eigenvectors of a matrix that is non-singular (determinant not = 0) are the values associated with linear transformations of that matrix. They are calculated using the correlation or covariance matrix functions. The eigenvalues are the values associated with the strength or the degree of a linear transformation (such as bending or rotating). See Linear Algebra by Gilbert Strang (online ebook) for more details on their computation.

4. How do the recommender systems in Amazon and Netflix work? (research paper pdf)

Recommender systems in Amazon and Netflix are considered top-secret and are usually described as black boxes. But their internal mechanism has been partially worked out by researchers. A recommender system, predated by expert systems models in the 90s, is used to generate rules or ‘explanations’ as to why a product might be more attractive to user X than user Y. Complex algorithms are used, which have many inputs, such as past history genre, to generate the following types of explanations: functional, intentional, scientific and causal. These explanations, which can also be called user-invoked, automatic or intelligent, are tuned by certain metrics such as user satisfaction, user rating, trust, reliability, effectiveness, persuasiveness etc. The exact algorithm still remains an industry secret, similar to the way that Google keeps the algorithms that perform PageRank secret and constantly updated (500-600 times a year in the case of Google).

5. What is the probability of an impossible event, a past event and what is the range of a probability value?

An impossible event E has P(E) = 0. Probabilities take on values only in the closed interval [0, 1]. The probability of event that is from the past is an event that has already occurred and here P(E) = 1.

6. How do we treat missing values in datasets?

A categorical missing value is given its default value. A continuous missing value is usually assigned using the normal distribution, or the measures of central tendency like mean, median and mode. If a feature has less than 20% available data, the recommendation is to delete that feature from the model.

7. Which is faster, Python or R?

Python is considered to be moderately medium-paced since C++ is much faster for all purposes. Besides which, Python is an interpreted and not a compiled language. Python language is implemented in C to speed up execution time. R, however, was designed by statisticians, not computer scientists, and is much slower than Python.

8. What is Deep Learning and why is it such a popular buzzword in the machine learning field right now?

For many years, until around 2006, backpropagation neural networks had just three layers – one input, one hidden and one output layer. The problem with this model was that since it used gradient descent and the backpropagation algorithm, the neural nets had a tendency to be attracted towards the local minima in the hyperplane that represented the dimensions of the input features. Thus, NNs could not be used for many applications optimally, since they could only find a partially optimal solution. In 2006, Geoffrey Hinton et. al. published a research paper that showed that multilayer neural networks could overcome the problem of local minima since, in thousands of dimensions, local minima are statistically so rare as to never be found in the back-propagation process (saddle points are common instead). Deep learning refers to neural nets with 3 or more (even 10) hidden layers. They require more computational power and were one of the reasons that GPUs started to be used by the machine learning community for implementation of deep learnings NNs. Since 2010-2012, deep learning has been applied to nearly every single technology domain, and the models have been highly accurate and successful in all areas from speech recognition to playing the Japanese game of Go.

9. What is the difference between machine learning and deep learning?

For more details on that, I suggest you go through this excellent article, given on the following link on our blog below:

https://dimensionless.in/machine-learning-and-deep-learning-differences/

10. What is Reinforcement Learning?

For an excellent explanation of reinforcement learning that is both educational and fun to read, please visit the following page, also on our blog :

https://dimensionless.in/reinforcement-learning-super-mario-alphago/

Enjoy Your Work!

To finally sum up, I have to say, enjoy your work. You will be much better at what you love than something that is glamorous but not to your taste. Artificial Intelligence, Data Science, Software Development and Machine Learning are very much in my preferred line of work, and my hope is, that it will be in yours too. Don’t just read the text, work out the code on your systems or on Kaggle. That is how to best prepare for interview questions. Only practice at your computer (preferably on Kaggle) will give you true confidence on the day of your interview. That is true expertise – practice making perfect. Enjoy data science!

1 Comment

Peter on February 28, 2019 at 4:17 pm

Checking out some of the links, it seems that TestDome also has Data Science specific questions: https://www.testdome.com/d/data-science-interview-questions/28

Though the Python questions are still good too. No R though.
Reply

Trackbacks/Pingbacks

Top 5 Ways to Evaluate Data Science Competency | DIMENSIONLESS TECHNOLOGIES PVT.LTD. - […] Data Science Interview Questions with Answers […]

Submit a Comment Cancel reply

Dimensionless Techademy

4.9

Based on 37 reviews

review us on

Dellima Stella

11:42 22 Nov 21

Never thought that online trading could be so helpful because of so many scammers online until I met Miss Judith... Philpot who changed my life and that of my family. I invested $1000 and got $7,000 Within a week. she is an expert and also proven to be trustworthy and reliable.
Contact her via:
Whatsapp: +17327126738
Email:judithphilpot220@gmail.comread more

Grace Leah

21:48 18 Nov 21

A very big thank you to you all sharing her good work as an expert in crypto and forex trade option. Thanks for... everything you have done for me, I trusted her and she delivered as promised.
Investing $500 and got a profit of $5,500 in 7 working days, with her great skill in mining and trading in my wallet.

judith Philpot company line:...
WhatsApp:+17327126738
Email:Judithphilpot220@gmail.comread more

Deepak Prasad

16:14 06 Apr 21

Faculty knowledge is good but they didn't cover most of the topics which was mentioned in curriculum during online... session. Instead they provided recorded session for those.read more

Ritika Khandelwal

09:06 14 Apr 20

Dimensionless is great place for you to begin exploring Data science under the guidance of experts. Both Himanshu and... Kushagra sir are excellent teachers as well as mentors,always available to help students and so are the HR and the faulty.Apart from the class timings as well, they have always made time to help and coach with any queries.I thank Dimensionless for helping me get a good starting point in Data science.read more

Rupal Gupta

07:33 27 Nov 19

My experience with the data science course at Dimensionless has been extremely positive. The course was effectively... structured . The instructors were passionate and attentive to all students at every live sessions. I could balance the missed live sessions with recorded ones. I have greatly enjoyed the class and would highly recommend it to my friends and peers.

Special thanks to the entire team for all the personal attention they provide to query of each and every student.read more

Durgesh Tiwari

11:26 05 Oct 19

It has been a great experience with Dimensionless . Especially from the support team , once you get enrolled , you... don't need to worry about anything , they keep updating each and everything. Teaching staffs are very supportive , even you don't know any thing you can ask without any hesitation and they are always ready to guide . Definitely it is a very good place to boost careerread more

Jasminder Singh

13:18 14 Sep 19

The training experience has been really good! Specially the support after training!! HR team is really good. They keep... you posted on all the openings regularly since the time you join the course!!
Overall a good experience!!read more

Akash Lamba

16:59 17 Aug 19

Dimensionless is the place where you can become a hero from zero in Data Science Field. I really would recommend to all... my fellow mates. The timings are proper, the teaching is awsome,the teachers are well my mentors now. All inclusive I would say that Kush Sir, Himanshu sir and Pranali Mam are the real backbones of Data Science Course who could teach you so well that even a person from non- Math background can learn it. The course material is the bonus of this course and also you will be getting the recordings of every session. I learnt a lot about data science and Now I find it easy because of these wonderful faculty who taught me. Also you will get the good placement assistance as well as resume bulding guidance from Venu Mam. I am glad that I joined dimensionless and also looking forward to start my journey in data science field. I want to thank Dimensionless because of their hard work and Presence it made it easy for me to restart my career. Thank you so much to all the Teachers in Dimensionless !read more

Harshal Marathe

13:15 17 Aug 19

Dimensionless has great teaching staff they not only cover each and every topic but makes sure that every student gets... the topic crystal clear. They never hesitate to repeat same topic and if someone is still confused on it then special doubt clearing sessions are organised. HR is constantly busy sending us new openings in multiple companies from fresher to Experienced. I would really thank all the dimensionless team for showing such support and consistency in every thing.read more

Shree Krishna Mishra

08:00 30 May 19

I had great learning experience with Dimensionless. I am suggesting Dimensionless because of its great mentors... specially Kushagra and Himanshu. they don't move to next topic without clearing the concept.read more

Jagdish Mishra

06:04 26 May 19

Dimensionless Machine learning with R and Python course is good course for learning for experience professionals.

Priyanka Gupta

06:10 29 Mar 19

My experience with Dimensionless has been very good. All the topics are very well taught and in-depth concepts are... covered. The best thing is that you can resolve your doubts quickly as its a live one on one teaching. The trainers are very friendly and make sure everyone's doubts are cleared. In fact, they have always happily helped me with my issues even though my course is completed.read more

Maulik J Patel

01:49 17 Feb 19

I would highly recommend dimensionless as course design & coaches start from basics and provide you with a real-life... case study.
Most important is efforts by all trainers to resolve every doubts and support helps make difficult topics easy..read more

Kaustubh Powar

12:35 15 Feb 19

Dimensionless is great platform to kick start your Data Science Studies. Even if you are not having programming skills... you will able to learn all the required skills in this class.All the faculties are well experienced which helped me alot. I would like to thanks Himanshu, Pranali , Kush for your great support. Thanks to Venu as well for sharing videos on timely basis...😊

Regards...

Kaustubhread more

Avneet Arora

08:50 15 Feb 19

I highly recommend dimensionless for data science training and I have also been completed my training in data science... with dimensionless. Dimensionless trainer have very good, highly skilled and excellent approach.
I will convey all the best for their good work.
Regards
Avneetread more

Jayakrushna Das

13:16 26 Jan 19

After a thinking a lot finally I joined here in Dimensionless for DataScience course. The instructors are experienced &... friendly in nature. They listen patiently & care for each & every students's doubts & clarify those with day-to-day life examples.
The course contents are good & the presentation skills are commendable. From a student's perspective they do not leave any concept untouched. The step by step approach of presenting is making a difficult concept easier. Both Himanshu & Kush are masters of presenting tough concepts as easy as possible. I would like to thank all instructors: Himanshu, Kush & Pranali.read more

Kiran Achanta

06:47 19 Jan 19

When I start thinking about to learn Data Science, I was trying to find a course which can me a solid understanding of... Statistics and the Math behind ML algorithms. Then I have come across Dimensionless, I had a demo and went through all my Q&A, course curriculum and it has given me enough confidence to get started. I have been taught statistics by Kush and ML from Himanshu, I can confidently say the kind of stuff they deliver is In depth and with ease of understanding!read more

Kumar Gaurav

15:23 08 Jan 19

If you love playing with data & looking for a career change in Data science field ,then Dimensionless is the best... platform . It was a wonderful learning experience at dimensionless. The course contents are very well structured which covers from very basics to hardcore . Sessions are very interactive & every doubts were taken care of. Both the instructors Himanshu & kushagra are highly skilled, experienced,very patient & tries to explain the underlying concept in depth with n number of examples. Solving a number of case studies from different domains provides hands-on experience & will boost your confidence. Last but not the least HR staff (Venu) is very supportive & also helps in building your CV according to prior experience and industry requirements.
I would love to be back here whenever i need any training in Data science further.read more

Rajesh Raj

17:14 25 Dec 18

It was great learning experience with statistical machine learning using R and python. I had taken courses from... Coursera in past but attention to details on each concept along with hands on during live meeting no one can beat the dimensionless team.read more

Deepak Singla

11:59 12 Nov 18

I would say power packed content on Data Science through R and Python. If you aspire to indulge in these newer... technologies, you have come at right place. The faculties have real life industry experience, IIT grads, uses new technologies to give you classroom like experience. The whole team is highly motivated and they go extra mile to make your journey easier.
I’m glad that I was introduced to this team one of my friends and I further highly recommend to all the aspiring Data Scientists.read more

Jitendra Yadav

10:58 27 Sep 18

It was an awesome experience while learning data science and machine learning concepts from dimensionless. The course... contents are very good and covers all the requirements for a data science course. Both the trainers Himanshu and Kushagra are excellent and pays personal attention to everyone in the session. thanks alot !!read more

Gunjeett Singh

18:09 10 May 18

Had a great experience with dimensionless.!!
I attended the Data science with R course, and to my finding this... course is very well structured and covers all concepts and theories that form the base to step into a data science career. Infact better than most of the MOOCs.
Excellent and dedicated faculties to guide you through the course and answer all your queries, and providing individual attention as much as possible.(which is really good).
Also weekly assignments and its discussion helps a lot in understanding the concepts.
Overall a great place to seek guidance and embark your journey towards data science.read more

Ajit Singh

08:31 21 Apr 18

Excellent study material and tutorials. The tutors knowledge of subjects are exceptional.
The most effective part... of curriculum was impressive teaching style especially that of Himanshu.
I would like to extend my thanks to Venu, who is very responsible in her jobread more

Yeshwanth Ram

11:33 01 Apr 18

It was a very good experience learning Data Science with Dimensionless. The classes were very interactive and every... query/doubts of students were taken care of. Course structure had been framed in a very structured manner. Both the trainers possess in-depth knowledge of data science dimain with excellent teaching skills. The case studies given are from different domains so that we get all round exposure to use analytics in various fields. One of the best thing was other support(HR) staff available 24/7 to listen and help.I recommend data Science course from Dimensionless.read more

Prabhakar Kumar

14:32 31 Mar 18

I was a part of 'Data Science using R' course. Overall experience was great and concepts of Machine Learning with R... were covered beautifully. The style of teaching of Himanshu and Kush was quite good and all topics were generally explained by giving some real world examples. The assignments and case studies were challenging and will give you exposure to the type of projects that Analytics companies actually work upon. Overall experience has been great and I would like to thank the entire Dimensionless team for helping me throughout this course. Best wishes for the future.read more

Megha Kansal

14:47 30 Mar 18

It was a great experience leaning data Science with Dimensionless .Online and interactive classes makes it easy to... learn inspite of busy schedule. Faculty were truly remarkable and support services to adhere queries and concerns were also very quick. Himanshu and Kush have tremendous knowledge of data science and have excellent teaching skills and are problem solving..Help in interviews preparations and Resume building...Overall a great learning platform. HR is excellent and very interactive. Everytime available over phone call, whatsapp, mails... Shares lots of job opportunities on the daily bases... guidance on resume building, interviews, jobs, companies!!!! They are just excellent!!!!! I would recommend everyone to learn Data science from Dimensionless only 😊read more

Jagdish Ahuja

07:30 05 Mar 18

Excellent teaching techniques.....
Both of them have a very unique and great grip of the subject ....

Saurabh Kandhvey

17:51 20 Dec 17

Nice people in terms of technical exposure .....very friendly and supportive. A place to start your Data Science... learning.read more

Ashwani Pandey

10:55 02 Dec 17

An awesome place to learn. Complete package of theritocal and practical knowledge.

Ashwini Ningdalli

11:08 31 Oct 17

Saroja Gundiga

09:45 31 Oct 17

I am very glad to be part of Dimensionless .Their dedication, in-depth knowledge, teaching and the way they explain to... clarify doubts is tremendous . I recommend this to everyone who wish to build their career in Data Science

With whole heartedly I wish them for their success & future prospectsread more

Ashish Mohan Sharma

11:47 17 Jun 17

Being a part of IT industry for nearly 10 years, I have come across many trainings, organized internally or externally,... but I never had the trainers like Dimensionless has provided. Their pure dedication and diligence really hard to find. The kind of knowledge they possess is imperative. Sometimes trainers do have knowledge but they lack in explaining them. Dimensionless Trainers can give you ‘N’ number of examples to explain each and every small topic, which shows their amazing teaching skills and In-Depth knowledge of the subject. Himanshu and Kush provides you the personal touch whenever you need. They always listen to your problems and try to resolve them devotionally.

I am glad to be a part of Dimensionless and will always come back whenever I need any specific training in Data Science. I recommend this to everyone who is looking for Data Science career as an alternative.

All the best guys, wish you all the success!!read more

12:13 11 Nov 16

09:51 13 Oct 16

12:14 29 Sep 16

07:05 02 Mar 16

Topics

Agent-Based Modelling (1)
AI (1)
Analytics (18)
Artificial Intelligence (2)
AWS (14)
Big Data (24)
- Learn big data (3)
Blockchain (3)
Blog (2)
Business Analysis (5)
Career Transitions (13)
Cloud Technologies (10)
Data Science (161)
- Learn Data Science (35)
- Testimonial (1)
Data Science Applications (12)
Data Science Cyber Crime (2)
Data Visualization (3)
Deep Learning (14)
Dream Job (7)
Future-Ready Careers (2)
Hadoop (1)
Interview Questions (1)
Julia (1)
machine learning (31)
Mistakes in Data Science (1)
Natural Language Processing (4)
NLP (4)
Projects (5)
Python (24)
Quantum Computing (2)
R Programming (12)
Scoop.it (7)
Statistics (5)
Training (10)
Trending (9)
Uncategorized (12)
Visualisation (9)