京公网安备 11010802034615号
经营许可证编号:京B2-20210330
By Matthew Mayo, KDnuggets.
It has been a year and a half since Linda Burtch of Burtch Works wrote 9 Must-Have Skills You Need to Become a Data Scientist, a post which outlined analytical, computer science, and non-technical skills required for success in data science, along with some resources for gaining and improving these skills. While this post is still relevant and quite popular, I thought I would take a shot at updating it, taking into account the direction of data science developments over the past 18 months.
My approach is a bit different than Linda's, which was to distill the views of, and conversations with, a number of analytics professionals considering adapting their skills to the field of data science at the time; mine is based on observations of trends, content of articles, prevalence of ideas, and discussions with a number of individuals in various positions of career development in the field. Please take this as additional information to take under advisement, as opposed to any kind of definitive advice.
Non-Technical Skills(非技术能力)
1. Education(教育)
Burtch provides some numbers related to the educational level of data scientists, indicating that 88% of Data Scientists have, at minimum, a Master's degree. Burtch does not explicitly provide her source, but I can only assume it comes from her firm's extensive research, with which I am not going to outright contradict. What I will offer, instead, is that data science is an incredibly diverse field, with no real consensus as to what it even is. I'm sure people will disagree with that, but when I hear the term data scientist, I tend to think of the unicorn, and all that it entails, and then remember that they don't exist, and that actual data scientists play many diverse roles in organizations, with varying levels of business, technical, interpersonal, communication, and domain skills. If we recall that different roles such as machine learning scientists, data analysts, data engineers, Hadoop administrators, and analytics-focused MBAs often get wrapped up in with the definition, it's easy to determine that there would be many paths to data science. To be fair, Burtch's definition of what a Data Scientist is, and therefore what educational levels they have, could vary quite drastically from mine.
That said, there is likely much more variety of education levels held by thoseconsidering themselves data scientists. To that end, the 2015 Stack Overflow Developer Survey, which is composed of self-reporting data, provides the following, for example:

Granted, this is a single snapshot in time, and it would be foolish to draw conclusions based on this single piece of data. However, the takeaway is further evidence to support the idea that there is no one "correct" path to data science; folks come from academia, industry, computer science, statistics, physics, other hard sciences, engineering, architecture... and they hold PhDs, Master's degrees, undergraduate degrees, and, yes, some are even *gasp* self-trained (probably). Don't let anyone dissuade you from pursuing "data science," and don't let anyone tell you you have to do X to become one. Find out the data science niche you want to fill, and pursue training and education that will allow for it. And be realistic: you may aspire to Chief Scientist at DeepMind, but self-learning with MOOCs and a few textbooks likely won't do it. But that doesn't mean that self-learning with MOOCs and a few textbooks won't get you somewhere interesting in data science. If the anecdotes are to be trusted, it happens regularly.
2. Intellectual Curiosity(求知欲)
This point requires far less fleshing out. Simply put: if you don't have intellectual curiosity, data science ain't for you. Next.
3. Domain Knowledge/Business Acumen(领域知识/商业智慧)
Whether we're talking theoretical unicorns or something closer to the data science professional periphery, you really need some understanding of the domain you are working in to be useful, analytically speaking. Consider even a purely technical role: if you are developing algorithms, pipelines, or workflows for an organization, without a solid understanding of the fundamentals of the industry and the goals of the firm, you won't be able to appropriately leverage your technical abilities to make a difference in the long run. And let's face it:making a difference is what data science is all about.
4. Communication Skills(沟通技巧)
Again, this isn't difficult to understand. Data science persons need real communicate good blah blah.
Burtch summed up the reasons for this in her previous iteration of the post: The "data scientist must enable the business to make decisions by arming them with quantified insights, in addition to understanding the needs of their non-technical colleagues in order to wrangle the data appropriately." Boom.
5. Career Mapping/Goals(职业映射/目标)
This is related to skill #1, education and training. Let's again think of the unicorn; he or she could, in all their glory, fulfill any analytical or executive position at any company within their domain knowledge range. The true unicorn (henceforth, the Trunicorn) would also be guaranteed immense sums of money. But chances are you aren't a unicorn, and never will be (are there anyunicorns???), and so you need to plan out your career path and execute on that plan.
If you want to maintain a more technical role permanently, plan to keep those technical skills in tip-top shape. If you want to crossover into a role that involves more interaction with clients, then brush up on your communication and presentation skills. If you think your niche is devising and administering big data processing solutions, then get your Apache on, dammit! Good things don't come to those who wait; good things come to those who devise detailed plans of action rooted in reality, and then execute on said plans.
Technical Skills(技术能力)

6. Coding Skills(编码技巧)
It's not "R vs Python" or "R and Python" or "R and Python vs something else." It's "what skills do I need to fulfill my duties?" As someone from a computer science background, these arguments (unfortunately) are par for the course. n00bs be like: "Java is waaay better than Assembly." Well, what is it you aim to do, and which language do you know better? And flame wars are dumb (especially this one).
R is very strong for pure data analysis. Python has a rich scientific ecosystem that better lends to development of software solutions and industrial-strength implementations. That doesn't tell the whole story, though, and there is lots of crossover. Data exploration in Python? Sure. Machine learning in R? Why not? If you know your tools well, you know when to use them. It's really that straightforward. Maybe you only need one, but maybe not. You need to know that as well.
And it doesn't stop at "R or Python?" All sorts of languages and libraries are useful for data science. Of note, Java and Scala have their place in big data processing, thanks to their prevalence in the ecosystems which grew up around the popular frameworks. A lot of low level coding is done in C++, especially for algorithm development, thanks to the speed and control associated with being closer to the metal. Tools are just that; they aren't meant to become ideological expression of dogma we associate with to form our identities. But you do need to be in possession of some tools.
7. Machine Learning/Data Mining Skills(机器学习/数据挖掘技能)
This refers to both theoretical and practical skills. You don't want someone with no idea of how kernel methods function or what higher dimensionality is to be implementing Support Vector Machines and hoping they can logically interpret results. At the same time, the demand for someone who could explain these concepts ad nauseam but not be able to implement an SVM classifier is probably quite low. And then, obviously, learning implementations based on particular environments would be required.
As a particular exemplar, see this post on mastering machine learning in Python, which starts with the theoretical and moves toward the practical.
8. Big Data Processing Platforms: Hadoop, Spark, Flink, etc.(大数据处理平台:Hadoop的,星火,弗林克等。)
See skill #6 for a discussion on not putting all your stock in any one technology or platform, but instead treating them like the tools that they are. Then see this post for an overview of contemporary big data processing frameworks. The point is this: data is growing, and as a data scientist you have to understand that data processing frameworks are a part of the data science landscape; having an understanding of these frameworks is vital.
9. Structured Data (SQL)(结构化数据(SQL))
Burtch was wise to point out, specifically, that there is a difference between structured and unstructured data skills, and that data scientists should (must?) be familiar with both. Structured data is synonymous with relational data, which is lorded over by the one query language to rule them all: SQL. There is abundant conflation of concepts here, but these days:
structured == relational == SQL
At the very least, it is expected that data scientists can write and execute non-trivial SQL scripts against stored data.
10. Unstructured Data (3-5 top NoSQL DBs)(非结构化数据(3-5顶部的NoSQL数据块))
There is much less... well, structure among the components of unstructured data storage and management. As such, different tools are required to store, retrieve, analyze, and otherwise process this data. The path to unstructured data storage and interaction is not as straightforward as it is for structured data, where relational database systems and SQL are the only real game in town. The NoSQL (I dislike the term, but it gets us where we need to go quickly) database, according to the internet's resident know-it-all, Wikipedia, "provides a mechanism for storage and retrieval of data which is modelled in means other than the tabular relations used in relational databases." Not very specific, but point taken.
Data scientists need to know how to manage unstructured data, and the options for doing so are many. Popular NoSQL database architectures include key-value stores, document stores, tuple stores, and wide column stores; each of these types have different approaches and philosophies, and the number of available implementations are seemingly endless (MongoDB, CouchDB, Cassandra, Druid, MemcacheDB...). The bottom line here is to know the terrain, study the architectures, and gain better-than-passing knowledge of one or two strong NoSQL database system implementations.
哈哈,如果看不下去英文文章,该文章的中文版会在接下来的公众号里呈现给大家,敬请期待!
数据分析咨询请扫描二维码
若不方便扫码,搜微信号:CDAshujufenxi
箱线图(Box Plot)作为数据分布可视化的核心工具,能清晰呈现数据的中位数、四分位数、异常值等关键统计特征,广泛应用于数据分 ...
2026-01-28在回归分析、机器学习建模等数据分析场景中,多重共线性是高频数据问题——当多个自变量间存在较强的线性关联时,会导致模型系数 ...
2026-01-28数据分析的价值落地,离不开科学方法的支撑。六种核心分析方法——描述性分析、诊断性分析、预测性分析、规范性分析、对比分析、 ...
2026-01-28在机器学习与数据分析领域,特征是连接数据与模型的核心载体,而特征重要性分析则是挖掘数据价值、优化模型性能、赋能业务决策的 ...
2026-01-27关联分析是数据挖掘领域中挖掘数据间潜在关联关系的经典方法,广泛应用于零售购物篮分析、电商推荐、用户行为路径挖掘等场景。而 ...
2026-01-27数据分析的基础范式,是支撑数据工作从“零散操作”走向“标准化落地”的核心方法论框架,它定义了数据分析的核心逻辑、流程与目 ...
2026-01-27在数据分析、后端开发、业务运维等工作中,SQL语句是操作数据库的核心工具。面对复杂的表结构、多表关联逻辑及灵活的查询需求, ...
2026-01-26支持向量机(SVM)作为机器学习中经典的分类算法,凭借其在小样本、高维数据场景下的优异泛化能力,被广泛应用于图像识别、文本 ...
2026-01-26在数字化浪潮下,数据分析已成为企业决策的核心支撑,而CDA数据分析师作为标准化、专业化的数据人才代表,正逐步成为连接数据资 ...
2026-01-26数据分析的核心价值在于用数据驱动决策,而指标作为数据的“载体”,其选取的合理性直接决定分析结果的有效性。选对指标能精准定 ...
2026-01-23在MySQL查询编写中,我们习惯按“SELECT → FROM → WHERE → ORDER BY”的语法顺序组织语句,直觉上认为代码顺序即执行顺序。但 ...
2026-01-23数字化转型已从企业“可选项”升级为“必答题”,其核心本质是通过数据驱动业务重构、流程优化与模式创新,实现从传统运营向智能 ...
2026-01-23CDA持证人已遍布在世界范围各行各业,包括世界500强企业、顶尖科技独角兽、大型金融机构、国企事业单位、国家行政机关等等,“CDA数据分析师”人才队伍遵守着CDA职业道德准则,发挥着专业技能,已成为支撑科技发展的核心力量。 ...
2026-01-22在数字化时代,企业积累的海量数据如同散落的珍珠,而数据模型就是串联这些珍珠的线——它并非简单的数据集合,而是对现实业务场 ...
2026-01-22在数字化运营场景中,用户每一次点击、浏览、交互都构成了行为轨迹,这些轨迹交织成海量的用户行为路径。但并非所有路径都具备业 ...
2026-01-22在数字化时代,企业数据资产的价值持续攀升,数据安全已从“合规底线”升级为“生存红线”。企业数据安全管理方法论以“战略引领 ...
2026-01-22在SQL数据分析与业务查询中,日期数据是高频处理对象——订单创建时间、用户注册日期、数据统计周期等场景,都需对日期进行格式 ...
2026-01-21在实际业务数据分析中,单一数据表往往无法满足需求——用户信息存储在用户表、消费记录在订单表、商品详情在商品表,想要挖掘“ ...
2026-01-21在数字化转型浪潮中,企业数据已从“辅助资源”升级为“核心资产”,而高效的数据管理则是释放数据价值的前提。企业数据管理方法 ...
2026-01-21在数字化商业环境中,数据已成为企业优化运营、抢占市场、规避风险的核心资产。但商业数据分析绝非“堆砌数据、生成报表”的简单 ...
2026-01-20