Does openness make an AI model more trustworthy?
Openness gives access to model materials; it does not certify the accuracy or safety of an application. The information available for inspection and the freedom to modify it depend on the release and its licence.
The Open Source Initiative distinguishes access to weights from a fuller set of code, data information and permissions. These materials can help a technical assessment, but their availability does not settle how a connected business system behaves. The OSI definition
Trust is also specific to the task. A model producing internal draft text presents a different exposure from a system publishing product claims or acting on customer records.
What matters about the model and its host?
The model’s identity, operating environment and access to information determine what is being assessed. A model card describes the released model; a hosting agreement describes a separate service around it.
OpenAI’s gpt-oss-20b card identifies its licence and provides deployment information. It does not certify another company’s hosted application merely because that application uses the model. The published model card
A business assessment can cover who operates the system, which records it receives and who can change its settings. Maintenance responsibilities also differ between a managed service and a model operated by the business itself. Neither arrangement removes the need for evidence about the actual use.
How do connected tools affect the risk?
Connected tools change what an AI system can read or do, so they can alter the consequences of an error. Permission to draft a message differs from permission to send it or spend money.
The NCSC describes prompt injection as a risk in which external material can interfere with a language model’s intended behaviour. Its December 2025 guidance concerns system design; it is not a performance comparison of today’s models. The NCSC explanation
The relevant assessment therefore includes the surrounding application and its permissions. Local hosting or access to weights does not, by itself, establish that those connections are safe.
What evidence supports a business assessment?
Task-specific evaluation provides evidence about whether a particular deployment works acceptably under defined conditions. It can include routine inputs, difficult cases and the consequences of incorrect outputs.
Results have a scope: the tested model version, available data and permitted actions. Changing those conditions can change what the evaluation establishes. An average accuracy score can also conceal a consequential failure in a small group of cases.
NIST’s voluntary AI Risk Management Framework provides organisational guidance for assessing and managing these questions. It is a management framework, rather than a certification for a named model. NIST’s implementation guidance
Legal note: This answer provides general information, not legal advice. Seek advice from qualified legal counsel for your circumstances.