Is it just a harmless mole, or is urgent action needed? The answer comes from looking at the computer screen: There, an artificial intelligence-based program highlights unusual skin changes that a dermatologist should examine more closely. Human medical expertise remains the foundation for the diagnosis. During the personal consultation, the further course of treatment is discussed – there has always been a special relationship of trust between doctor and patient. But what happens in such a case when AI becomes increasingly involved in the diagnosis? “There is trust between doctor and patient, but there is also trust in the technology being used,” says Joachim Bürkle, Managing Director of the AI Quality & Testing Hub (AIQ), in which VDE and the state of Hesse are shareholders and which develops reliable quality testing for AI systems.
The example from the dermatology practice shows how close AI has become to the everyday lives of many people. The technology has long been operating “invisibly” in critical infrastructure, vehicles, or assistance systems—and therefore in areas where safety is essential. “Trust in AI is crucial because AI has now entered every area of our lives,” says Joachim Bürkle.
But trust is not always the same kind of trust: Martin Schneider, Group Leader for Testing at the Fraunhofer Institute for Open Communication Systems FOKUS, distinguishes between the trustworthiness of AI and the trust that users place in these technologies. Trustworthiness depends primarily on various quality parameters, including robustness, safety, transparency, and explainability. “Explainability is very important for trustworthiness,” he explains, referring to the buzzword “Explainable AI.” This refers to approaches that make AI decision-making processes understandable, verifiable, and controllable. In the past, programs operated largely deterministically. With AI, by contrast, the so-called “black box” is frequently criticized, not least because of errors or hallucinations.
There are various approaches to shed light on the black box. Among the promising ones are the LIME and SHAP methods. Put simply, LIME – the abbreviation stands for “Local Interpretable Model Agnostic Explanations” – shows which factors influenced the result. For example, in our dermatology practice, LIME would show which pixels in the camera image were responsible for the area being flagged as suspicious. The “Shapley Additive Explanations” model, SHAP, goes a step further and calculates how much a particular piece of information contributed to an AI decision. “You vary the input to find out exactly what might have played a role,” says Martin Schneider. Such approaches create the technical prerequisites for making decision-making processes more understandable – and thus form the basis for the regulations of the EU AI Act.
With the regulation, the European Union has taken a pioneering role. It establishes rules for the use of AI and takes a four-tier, risk-based approach. The higher the risk, the stricter the rules. “In high-risk areas, the issue of trust plays a greater role,” explains Joachim Bürkle. The AI Act defines the use of systems in areas such as critical infrastructure, education, and the justice system as high-risk.
Putting AI in the sandbox – but doing it right
How quickly trust in AI can reach its limits without these regulations was demonstrated by the cyberattack carried out by an OpenAI model on the Hugging Face platform in July. Assuming that the relevant data needed to solve the industry test ExploitGym could be found on Hugging Face, the model gained access to the company’s data. Normally, such AI models are tested using so-called “sandboxing” in an isolated environment. According to expert Schneider, the AI managed to break out of the sandbox. He sees this as a failure on the part of those responsible: “There is a weakness in the framework surrounding the AI model.” The AI should have been taught what it was allowed to do – and what it was not allowed to do.
In addition to sandboxing, the researcher recommends using so-called “guardrails.” “This limits the degrees of freedom available to the model.” That is precisely what did not happen sufficiently in the Hugging Face attack: OpenAI is said to have deliberately reduced the safeguards.