top of page

IFDS series part 2: Knowledge discovery reflections

5 days ago
6 min read

A series of blogs about the progress towards the Internet of FAIR data and services.



We also need a knowledge infrastructure


Making data FAIR is never a goal in itself, it is needed to enable continued and cumulative knowledge discovery and actual societal innovations. Moving from ‘Reusable’ data to data that are actually ‘Reused’ we need supporting infrastructure. As actual innovation; turning new knowledge into policy, as well as practical, scalable and sustainable applications, is not the core business of science, here we will focus on the process of the generation and communication of the knowledge itself. Knowledge, and thus also science informs policy. The International Science Council recently urged the next UN secretary-general to treat science as a core pillar of multilateral governance, warning that scientific knowledge remains “too often an ad hoc input to decision-making” even as governments confront increasingly complex challenges ranging from climate change and inequality to artificial intelligence and biotechnology."


This urge is in fact calling for a ‘Living Evidence Engine’ of which the fuel is trustworthy data of different kinds. The resource should be equally available to scientists, policy makers, industry and any other relevant actors, also in resource poor areas, as the SDG challenges are strongly felt in these regions of our planet. In the report of the European Open Science Cloud (EOSC) design phase this vision was already coined under the term the ‘Internet of FAIR data and services’ (IFDS).


Exchange knowledge not data


To really understand what we are talking about here, it is useful to briefly try to fundamentally understand how we actually discover knowledge. In the figure below, we summarise a model we have used internally for years to try and understand the basics of knowledge discovery in complex, multimodal biological data. What the model made visible is that if our ‘ability to understand’ (pattern recognition and translation to insights) would have no upper limit, we essentially continue to see all possible patterns in data. The line goes straight till infinite complexity. However, our human ability to discern and interpret complex patterns does have an upper limit. At that level, here depicted at an arbitrary cutoff as a horizontal dotted line, in case the data and the implicit patterns in them become even a little bit more complex, we (as humans) suddenly are ‘completely lost’ and perceive almost everything beyond that point of complexity as ‘chaos’. Unfortunately, we cannot connect millions of human minds, like we can combine CPUs and GPUs. Thus powerful computers, equipped with well trained models, will always easily outperform humans in pattern recognition in high-dimensional complex data sets. Dependent obviously on what we exactly define as ‘intelligence’ I would argue that this pattern recognition, is not comparable to human intelligence, where these patterns (even if ambiguous and blurry) may lead to new insights.



The model we used to create machine assisted hypotheses based on all captured biological knowledge in literature and databases.


Around the ‘percolation point’ (where order turns into perceived chaos), is where humans, increasingly with the assistance of computers) discover new patterns, insights and eventually knowledge.


When we deal with sensitive data, about which I will write more extensively later, it is much safer to share knowledge than to share data and information.


In the days when we used this model for biological hypothesis generation on a daily basis (way before LLMS bursted on the scene), we used the term ‘solid’ for what we now coined ‘established knowledge' (EK). This is more or less ‘fixed’ knowledge that is published and commonplace, and on which most ontological and conceptual models of our world have been built. In the ‘gas’ phase, we considered many of the patterns that computers seemed to discern as spurious as the concepts seemed to be in ‘brownian movement’. We now think that that phase is clearly associated with the ‘hallucination’ perception. Even if the correlations revealed by computers may even make some sense, they do not lead to any ‘actionable knowledge’ for people to act on. It is in the ‘fluid’ phase that we consider associations to be of high interest for further exploration. As the picture indicates, we considered most of multi-omics biology still out of range for human understanding. If people are interested we can provide the scientific articles that resulted from this research.


Policy and Discovery


Please note in the context of this article, that some of the patterns machines may discern in highly complex data may actually be revealing potentially actionable knowledge, and may be ‘mistaken’, at least initially, by humans as ‘hallucinations’. This becomes especially relevant when we want to use knowledge to inform policy and interventions, as we have to work mostly cross-disciplinary, with potentially less than optimal established knowledge in the collective intelligence of the users. Therefore, it is extremely important to be able to constrain the output of any ‘evidence engine’ to meaningful, trustworthy and ready to use information for the sharing and creation of relevant actionable knowledge. As long as the model of choice is constrained by conceptual boundaries that prevent machines from spitting out impossible assumptions, we should take all ‘predictions’ seriously. In the study of complex disease/genome associations we have used pattern recognition models for several decades and the model summarised here convinced us that actual knowledge discovery indeed takes place at ‘the edge of chaos’. In the scope of this article, it is crucial to acknowledge, that even if high dimensional data are analysed by skilled domain experts and their top ‘informaticians’ it is very difficult to distinguish golden nuggets of new meaningful causal relationships from spurious correlations.


Now that we translate this model to the concept of a ‘Living Evidence Engine’ where in many cases complex data are dispersed and multidisciplinary in nature, we need to optimise the trust in the outcomes, also now referred to as XAI, explainable AI. It is therefore crucial to use EK and approved conceptual (domain) models to constrain AI technology and its outputs. In research, in fact, knowledge discovery will change our conceptual model of the world.


In other words, the line representing the ‘percolation point’ in the model is consistently moving to the right. Knowledge discovery, the core business of science, is the progressive deeper understanding of reproducible patterns in what we perceive as ‘reality’. The more we discover, the more the model predicts vast ‘implicit knowledge’. The general perception that the ‘implictome', tacit and implicitly present meaningful patterns, will grow one order of magnitude faster than the ‘explicitome’, everything we already explicitly stated as factual knowledge.


The distinction between ‘intervention’, implementation and the scientific discovery process is particularly relevant in the coming decades because, as explained in the first blog post, we have reached the point where the data relevant for the societal and planetary problems we need to address today are so complex that computers see many patterns in them, but humans are no longer able to discern those without the help of these machines. But, as said before, machines just discover patterns but ‘have no clue’ what they mean. In other words, the closer we are to implementation the more ‘solid’ we want our outputs to be, while science will mostly be interested in the ‘fluid’ phase. One of the major skills of scientists (and policy makers) of the future is to fine-tune the ‘conceptual model filter’ in such a way that for societal interventions only ‘established knowledge’ and ‘clear trends’ are informing the policies and interventions, while allowing more degrees of freedom when it comes to a more exploratory and speculative phase, where the need for ‘hybrid intelligence’ and human expert groups may be crucial. This is for instance what happened in many countries during the COVID19 pandemic when new data and information were flooding the system so fast that the ‘percolation point line was moving much faster than in ‘normal periods’


Now, there - again - is also good news: we as humans are getting better and better at letting machines do this hard work of revealing patterns in massive amounts of data and also output it in human readable formats. We then compare what we already know (established knowledge) with real world observations. For humans the challenge remains to unravel the complex patterns and create actionable knowledge in our reality one little step at a time.


In the current ‘AI’ hyped discussions around LLMs it all comes back once more: We slowly realise that unambiguous, machine-actionable FAIR Digital Objects (FDO), accumulated in more and more complex FAIR objects that machines can operate on, will massively reduce meaningless hallucinations. In addition, we can restrict machines at the ‘output level’ with so-called ‘conceptual models’ that prevent meaningless outputs. Finally, it remains crucially important to not only endow FAIR data with rich provenance, analytical and governance metadata to allow machines to seamlessly operate on them, but to also describe the data in human readable formats, including the metadata regulating and enabling reuse (preferably visiting) by others, but also with the human readable description of the workflows used, the reasons for generating the data and so forth….


The TRYANGLE concept as the basis for a ‘living evidence engine’ to be elaborated in the next blog post will go deeper into these aspects.





 
 
 

Comments


bottom of page