Small Data, Big Value
Author: Mason Dykstra, Ph.D., VP Energy Solutions
Enthought welcomes Mason as VP, Energy Solutions, whose background from Anadarko, Statoil, and as a professor at the Colorado School of Mines qualifies him to make the case for ensuring ‘Small Data’ is equally part of the Fourth Industrial Revolution. The first in a Small Data series.
The origin of the term Big Data will likely never be agreed. However, in the world of science and computing, the case can be made that the term originated in Silicon Graphics in the 1990’s, whose work in video, for surveillance and Hollywood special effects, had it facing orders of magnitude more data than ever before. Recent advances in scientific computing technology and techniques, and massive generation of data, in particular by consumers and from social media, have put the term Big Data at center stage.
However, in many scientific fields, Big Data does not exist. It’s all about getting the most from ‘Small Data’, ensuring scientific challenges with minimal data also benefit from the ‘Fourth Industrial Revolution’. In many natural sciences and engineering disciplines large volumes of data can be hard or very expensive to generate. The reality is these datasets are often limited in size, poorly curated, and bespoke to particular problems. So, either the fields lacking in Big Data will be left out of the ‘Revolution’, or we need to work on ways of unleashing the power of Small Data.
Scientists are particularly adept at teasing meaning out of Small Data and drawing important conclusions with limited datasets. The future will be a collaboration between humans and machines, but clearly we don’t only want to solve the problems that have Big Data behind them. In cases where datasets are relatively small, or important pieces of information are missing, how can we develop this type of ‘intelligence’ in machines?
We need to engineer applications that can approach problems the way a scientist would. Scientists typically hypothesize as they go, which is to say they don’t wait until they have enough data to draw conclusions, but they actually generate, evolve and discard hypotheses along the way. While gathering data we are already engaging in problem-solving.
For example, when a geologist is creating a map of the geologic layers and faults under the Earth, they continually make educated guesses about what some of the map features will look like before they have gathered all the data. Not only does this give the geologist something early on paper (ok, on screen), but actually it provides a basis for hypothesis testing, and can help steer the succeeding data-gathering step. Think of this as akin to coming into a new town for the first time – even though you might never have been to that particular town before, all towns share certain traits and tend to have similarities which we can use to imagine the parts we haven’t yet seen. This kind of intuitive thinking and rule-of-thumb-based guessing, although critical for many sciences, has not been the realm of computers. Yet.
So the real question is can we capture the essential parts of that rule-making process and combine it with ‘machine reasoning’ to develop Small Data approaches that are akin to the way a scientist would approach a problem? But much faster and more consistent? This is one of the major challenges for many scientists today, whether they recognize it yet or not.
One thing we do know, paraphrasing Antonio di Leva in The Lancet; ‘Machines will not replace scientists, but scientists using AI will soon replace those not using it.’
About the Author
Mason Dykstra, Ph.D., VP Energy Solutions at Enthought, holds a PhD from the University of California Santa Barbara, an MS from the University of Colorado Boulder, and a BS from Northern Arizona University, all in the Geosciences. Mason has worked in Oil and Gas exploration, development, and production for over twenty years, split between oil industry-focused applied research at Colorado School of Mines and the University of California, Santa Barbara; and within companies including Anadarko Petroleum Corporation and Statoil (Equinor).
Related Content
イベントレポート | DXからインテリジェンス・ドリブンへ ~
エンソート創立25周年を記念した特別版として開催された本年度のサミットは、R&D領域のリーダーや経営層が一堂に会し、150名以上にご参加いただきました。 「DXからインテリジェンス・ドリブンへ ~ 科学的R&Dイノベーションの次代」をテーマとし、AIが加速する時代において研究開発組織がいかに進化すべきかが多角的に考察。本記事では各セッションの要点をレポートいたします。
科学的R&Dの転換点 ── エージェンティックAIは何を変えたのか
最先端AIを開発する企業は、いま、次々と科学研究の領域へ本格的に参入しています。これまで解決が難しかった科学的課題に対し、AIを「単なる支援ツール」ではなく、研究開発プロセスそのものを前進させる協働パートナーとして活用する時代が始まっています。本記事では、AIが「個別タスクを支援する技術」から、「研究開発プロセス全体を横断的に支援し推進する協働者」へと進化し始めた背景を具体的に紐解きます。
コンカレント材料設計:AIで実現する次世代アプローチ
AIの高度最適化、生成AI、エージェント型AIの活用により、材料と製品を同時に設計・最適化するコンカレント材料設計についてご紹介します。開発スピードと自由度が飛躍的に向上させることで、性能向上や市場投入までの期間短縮、競争優位性の確立が可能となります。
「収益性の壁」を超える:AIの活用で機能性材料開発を戦略から再構築
スペシャルティケミカルおよび素材産業は、コモディティ化と価格競争の激化により、従来の差別化戦略では持続的成長が難しくなっています。こうした中、AIやマテリアルズインフォマティクス(MI)などの先端技術が、R&D戦略の再構築と成長再加速への有力な打ち手として注目されています。
研究開発組織の変革を成功させるためのパートナー選び
現在の競争が激しいR&D環境において、適切なテクノロジーパートナーを選ぶことは、組織にとって最も重要な意思決定の1つです。理想的なパートナーとは、単なるツールベンダーやシステムインテグレーターではなく、生産性を向上させ、イノベーションを加速し、競争力を引き出す解決策を提供する科学的な専門知識と戦略的な洞察を兼ね備えた「変革の同志」です。
「AIスーパー・モデル」が材料研究開発を革新する
近年、計算能力と人工知能の進化により、材料科学や化学の研究・製品開発に変革がもたらされています。エンソートは常に最先端のツールを探求しており、研究開発の新たなステージに引き上げる可能性を持つマテリアルズインフォマティクス(MI)分野での新技術を注視しています。
デジタルトランスフォーメーション vs. デジタルエンハンスメント: 研究開発における技術イニシアティブのフレームワーク
生成AIの登場により、研究開発の方法が革新され、前例のない速さで新しい科学的発見が生まれる時代が到来しました。研究開発におけるデジタル技術の導入は、競争力を向上させることが証明されており、企業が従来のシステムやプロセスに固執することはリスクとなります。デジタルトランスフォーメーションは、科学主導の企業にとってもはや避けられない取り組みです。
産業用の材料と化学研究開発におけるLLMの活用
大規模言語モデル(LLM)は、すべての材料および化学研究開発組織の技術ソリューションセットに含むべき魅力的なツールであり、変革をもたらす可能性を秘めています。
科学研究開発における効率の重要性
今日、新しい発見や技術が生まれるスピードは驚くほど速くなっており、市場での独占期間が大幅に短縮されています。企業は互いに競争するだけでなく、時間との戦いにも直面しており、新しいイノベーションを最初に発見し、特許を取得し、市場に出すためにしのぎを削っています。