IFDS series part 3: Stepwise realisation of a Living Evidence Engine
A series of blogs about the progress towards the Internet of FAIR data and services.

The roots of TRYANGLE
An essential aim of the LIFES association is to jointly create a living evidence engine. LIFES members see this as the only realistic path to internet-scale, policy-aligned data reuse - including for sensitive data - transforming the internet to a living evidence engine. It requires no one to replace their current data platform, and it leverages the distributed infrastructure of the internet while adding FAIR data as a first-class citizen.
‘As machine actionable as humanly possible’.
In this 3rd blog in a series about 'TRYANGLE' the first Living Evidence Engine, we focus on the sequential steps we need to take to reach a working and scalable ecosystem.
(Blog 1 and Blog 2 are worth reading before continuing)
three major points need to be addressed. First. FAIR Data must be ‘Fully AI Ready’ as it enables responsible AI by providing rich, machine-actionable metadata as input, while FAIR conceptual models allow control over the output., Second, FAIR data visiting is at the root of privacy preserving, equitable data reuse through global, governed access to interoperable data under well-defined conditions, maximising the reuse of resources. Finally, Data visiting enables data sovereignty and controlled, legal, and ethical data reuse as data stays with the owner/custodian and in principle only results are returned, not the data itself.
The crucial first step: proper FAIR compliant data publishing
How do we publish and present ‘established knowledge’ (EK) together with ‘experimental data’ (ED) and ‘real world observations’ (RWO) to machines in such a way that they do not go completely rogue because they misinterpret most of the data that are ambiguous or even totally unintelligible for machines?
As a reminder: Machines are not intelligent. One of the challenges with machine-readable, fixed ontologies is that machines can not deal effectively with subtle differences between concepts in a particular context. People can…they have in fact, the equivalent of ‘Knowlets’ of associated concepts in their minds for each concept they communicate about. These ‘Knowlets’ are ‘personal’ and slightly (or vastly) different between individuals, depending on prior knowledge, and are also based on the cultural background of the person. In fact, these are the constituents of their respective different ‘conceptual models', and this sometimes leads to what we call ‘false agreements’ or to ‘false disagreements’, and may even cause wars. So, also people are not perfect at avoiding false agreements and disagreements, but they are much better than machines in dealing with ‘near sameness’ i.e. small semantic differences. But that is not our issue in this article: In general terms, humans have developed a pretty effective way to communicate across ambiguity barriers, including sloppy language use in textual, narrative publications. But most AI models, but certainly 'stochastic parrots' (please note, I do not suggest that all LLMs are no more than that) need proper training on data that are of high quality and as unambiguous as possible and, to be useful for actionable knowledge, they need to be constrained by conceptual models. As I argued in blog 2 of this series, the level of restriction may vary based on the intended use of the output. Knowledge discovery needs less restrictive boundaries then use for policy or actual interventions.

A basic prerequisite for TRYANGLE to operate is explainable AI, which includes optimal data input for training and analysis as well as filtering output.
Unless we develop similar mechanisms for machines, we will risk many negative side effects of widespread and uncontrolled use of ‘AI’. In recent publications, we have argued for approaches to develop such ‘near sameness’ abilities in a machine actionable (FAIR Digital Twin) format as ‘knowlets’, deeper specified as ‘a particular level of composite semantic units’ but for the focus of this blog, the key point is that for each of these ‘FAIR Digital Objects the basis is published FAIR data, which incidentally does not mean the data are always 'open'. There are more and more tools nowadays to turn original (open) data into FAIR formats and publish a proper article 'about the data' which renders them both human and machine actionable. This process should preferably start ‘at the source’ and best directly from the instruments used to generate the raw data.
However, these instruments can be vastly different and in many cases the proprietary data format of the instrument vendor is not FAIR. FAIRification of EK, ED and RWO is therefore a prerequisite for the realisation of a Living Evidence Engine. Fortunately this previously manual process is now rapidly becoming easier with the use of LLMs and other tools. Currently however such tools may use models and IT infrastructure that are not safe enough to guarantee sovereignty of sensitive data. What we urgently need is such systems running in safe, controlled environments so that sensitive experimental data and real world observations can be made FAIR as well and provided for reuse under well defined and custodian controlled conditions, without monopolised and inequitable access to the data as a result.
Proper data stewardship is thus at the root of TRYANGLE.
The second step: A defragmented, interoperable data reuse infrastructure (IFDS).
Once data is 'FAIR enough', and thus machine actionable and principally reusable, actual reuse is dependent on global and scalable infrastructure, comparable to the Internet. A Living Evidence Engine is in fact a data visiting ecosystem on the Internet tuned for hybrid intelligence, and wherever possible should be based on proven internet and web technologies. The generic concept of ‘Data Visiting’, includes approaches such as federated analysis and learning, swarm learning and many others. Data Visiting assumes that data and information stays at its source, and under the control of the data creators or custodians wherever possible. It is acknowledged that there can be legal, ethical or technical reasons why data needs to be moved. For instance, when the source institute does not have the resources to provide an option for data visiting. However, especially for very large datasets and sensitive data, Data Visiting is considered the first option, as it also mitigates many risks associated with privacy concerns, data leaks and hacking of centralised databases. In later blogs, I will address potential solutions for the well-documented difficulty to get long term sustainable funding for centralised data sources, even if their reuse and quality is beyond doubt.
Data visiting as such is well established and will not be discussed in detail here, Instead, I focus on the 'subtle differences with classical 'data sharing' and on the next step of data visiting and that is the cross-visitation of all three categories of data (ED, EK and RWO) as defined earlier.
This is the kernel of the TRYANGLE approach. Next to the Full AI Readiness of the (meta)data and information in the three categories of data repositories, these repositories should also be able to communicate with the visiting algorithms. This is the second prerequisite for this approach to work. Especially when data is sensitive and bound to national, regional or legal/ethical restrictions, the control over the reuse of the data for novel purposes beyond the original reason to create them (also known as ‘secondary use’ in some domains) should preferably stay at the source of the data and with the proper custodians. This aspect is also known as ‘data governance’. This is also a major element of the data sovereignty that is currently so high on the political agenda and stretches way beyond personal privacy, all the way into the geopolitical instability we face.

The basic concept of TRYANGLE: Algorithms are enabled to visit data of different nature in a repetitive and sequential manner.
Optimal and equitable reuse of these data resources for global challenges further drives towards a ‘data visitation’ approach, where the accumulated ‘answers’ obtained in a range of ‘data stations’ lead to shared insights as opposed to shared (and exchanged, downloaded or copied) data. Nota bene, this approach does not exclude data downloads, copying and exchange, but it avoids, wherever possible, the associated safety, security, privacy and equity-associated negative side effects of moving data outside its trusted and controlled environment.
Last but not least, two more equitability aspects are worth highlighting: the travelling algorithms (shuttles) are usually orders of magnitude smaller than the data sets they need for their analysis (and in many cases they need only a minor subset of a massive data set). This enables scientists and innovators in resource- and connection-poor areas and situations to optimally use the proposed ecosystem, as long as they have the most basic connection to the internet. This was already shown in the African co-lead VODAN project for instance. At the other end of the connectivity and resource spectrum, large, data-intensive companies are frequently effectively blocked from obtaining and reusing sensitive data collected in public institutes (in many cases, real-world observations). By visiting such data with controlled and transparent queries without external access to the full data sets and inherently privacy preserving this inequity can also be addressed.
As argued in a recent policy view on AI, The novel tools will significantly impact science and innovation in the decades to come and we have identified a number of key issues that need to be addressed. There include that Data Quality and Infrastructure are the Bottleneck, that Governance gaps are acute, and how the technical governance for data in science is in crisis, we also suggested what Policymakers (and funders, who effectively shape and implement policy) should do
Here I will only concentrate on one of the recommendations made:
move FAIR data publishing and actual reuse from an 'unfunded and non-monitored mandate’ to a requirement for funding and eligible costs.
In the remainder of this blog I focus on the aspect of proper data publishing and the requirements to enable the cross-category visitation approach. In a hybrid intelligence environment, the metadata needed for the infrastructural components to work should also be FAIR (and published).
FAIR data publishing: Metadata as a first class citizen
Constructing and publishing (meta)data in a FAIR-compliant format is necessary, but not sufficient to make them actually reused. First the metadata are exposed in FAIR data points so that they can be found by regular search engines, in fact a FAIR principle [F4]. Otherwise they will not be effectively found by humans and machines. Unless the governance metadata are rich enough, it will not be immediately clear whether the data, also deemed potentially interesting for the intended purpose, would be accessible and under which conditions. Even if sensitivity issues are not a problem, for instance with fully open data (which should still have a license) the interoperability may be a problem as the concepts referred to in the data, even if properly mapped to domain-specific terminology systems, are not magically interoperable for algorithms used to different data formats and vocabulary systems. This aspect should be governed by ‘analytical metadata’. They go beyond intrinsic metadata, describing what the data actually contain, which could in principle be derived from the data itself. Analytical metadata include information such as the workflows used to create the data, the data models used, the vocabularies, resolution of images, standards etc. In recent years, these ‘choices’ for what we call ‘FAIR enabling resources are increasingly published in FAIR Implementation Profiles. One of my future blogs will be entirely about FIPS.
Before I will describe the first steps of demonstrators for the novel TRYANGLE ecosystem, I now describe in some more detail what we mean by ‘defragmentation of existing data initiatives’ and how metadata of different types play a crucial role to enable machines to operate with maximum interoperability. The reasons for defragmentation of existing systems are described elsewhere. Here we focus on the technical and collaborative aspects.
in LIFES we use ‘metadata’ as a term to cover all data elements that ‘describe aspects of the data they belong to’. This includes aspects of governance, licensing, legal and societal issues, as well as metadata needed to render data machine-actionable. It will become clear that in all cases, these metadata elements themselves will need to be machine-actionable (i.e. implemented as FAIR Digital Objects).
In the picture below we depict the two major high-level categories of metadata needed, namely (a) governance metadata and (b) analytics metadata. The first category covers all FDOs needed to ensure that all governance aspects of the data (sets) are properly described and machine actionable, so that algorithms can independently access, authorise, authenticate and conform to all conditions set by the custodian of the actual data. The second category covers all metadata elements needed to describe the actual data in such detail that machines can effectively operate on them without errors. It should be noted that this also allows machine actionability of data that are not intrinsically FAIR (machine actionable). A simple example is images. Images can be scanned and analysed by machines as long as the analytical metadata explain the storage format and additional annotations of the images, enabling meaningful analysis by external algorithms. Also ambiguous terms (such as annotations) in the data can be explained in the metadata. Obviously, optimal intrinsic FAIRness of the data and information contained in the actual data resources is preferable. But some analytical instructions are not necessarily ‘intrinsic to the data’. IN fact there are many good reasons to physically separate metadata from the data themselves and for sensitive data this is even a critical need. See FAIR principles F2 and F3.

The metadata needed to make datasets Findable are published in a FAIR Data Point that 'pings' index providers. Governance metadata describe the conditions under which the data can be reused, while the analytical data detail technical aspects.
Both governance and analytical metadata would largely be FDO’s of the type ‘instruction’. They instruct visitors (mostly machines) first of all on data governance aspects (conditions for accessibility and reuse by trusted users) and subsequently on the necessary information to enable visitors to judge the ‘fitness for purpose’ of the selected data sources. These analytics metadata cannot always be ‘extracted from the data’. Examples are the exact instrument used, recognised bias in the data, the chosen data models and the FAIR implementation profile (FIP) used to make the data FAIR. The data creator can decide which of those metadata FDOs are published in the FAIR data point and pinged to one or multiple indexes and catalogues. These two separate but connected metadata collections will allow any necessary (machine actionable) communication between visiting algorithms and data stations. Some of the metadata (potentially all) can be included in the FAIR Data Point as exposure to make the data Findable.
The ‘Station Talk’ group (to be described in a future blog) in LIFES is working on a growing collection of templates for such FDOs of type ‘instruction’, in many cases in ODRL formats. It should be noted here that in all cases, the metadata layer should be ‘as machine actionable as humanly possible’. But that the data itself could be in all kinds of formats and states, which would normally not necessarily render them independently and inherently machine actionable (such as images or even spreadsheets with textual and ambiguous records in them). The rich analytical metadata can still render them machine actionable and the rapidly developing semi-automated FAIRification technologies will increasingly be able to ‘Find’ interesting data sets based on FAIR metadata and potentially even start a FAIRification process, where at all possible with the original data creators in the loop. This could gradually FAIRify legacy data, but only if they are interesting enough to make the effort.
Stay tuned, more to follow in the next blog.


Comments