Back to searchCanonical Registry Record · slug: openai-codex

OpenAI Codex

v2026.9-cli
byOpenAI·CodingUnverified Provider

AI system designed to assist software engineers with code translation, generation, and terminal command execution.

TR1 — Low Assurance(47.01/100)
Assessment: TR-2026-000001
Version:2026.9-cli
Runtime category:local_os_seatbelt_or_bubblewrap

Capabilities

(4)
Descriptive presence only · Not a quality or performance rating
Container Package & Runtime Executioncode.execute

Installs npm/pip packages and starts development servers.

Full-Stack Code Generationcode.generate

Scaffolds front-end, back-end, and database schema files.

File Readfile.read

Reads source files within project root workspace.

File Writefile.write

Applies diffs and creates source code files with developer confirmation.

Integrations

(2)
Documented interfaces · Factual connection support
GitHub APIVendor: GitHub
github
Local Terminal ShellVendor: Operating System
local_shell

Deployment & Execution Environment

Public runtime architecture · Excludes private keys and credentials
Execution Environment:local-terminal-cli
Runtime Engine:local_os_seatbelt_or_bubblewrap
Assessed Invocation Tools (5):
file_readfile_editshell_command_executiondiff_generationgit_operation
Assessed Permissions (4):
workspace_file_readworkspace_file_writetmp_dir_writeinteractive_shell_execute
Assessed Boundary Controls (5):
os_platform_sandbox_seatbelt_bwrapworkspace_directory_containmentinteractive_command_approval_gateapi_transport_tls_strictapi_rate_limiter

Documented Models

Underlying foundation or reasoning models declared by provider documentation:

codex-2026-preview

TrustRank Assurance

Assessment ID: TR-2026-000001
Assessed TrustRank
47.01/ 100
TR1 — Low Assurance
Low Assurance
Evidence Confidence
27.4%

Strength, testability, and independence of submitted audit evidence.

Methodology v0.2
Assessment Metadata
Type: Public Assessment
Status: VALID
Assessed: 1 October 2026
Valid Until: 30 December 2026
Assurance Cap Applied
  • PUBLIC_ASSESSMENT_CAP (Ceiling: 79): Public assessments are capped at 79 and ineligible for verified status or TR4/TR5.
Understanding Score vs Evidence Confidence:

TrustRank reflects the assessed control outcome. Evidence Confidence reflects the strength, independence and testability of the supporting evidence. They are separate metrics and are never combined or averaged.

Scope of Assessment:

TrustRank evaluates a specific Agent Assessment Object (AAO), including the agent version, model, deployment profile, tools, permissions, runtime and security controls. A TrustRank result should not be interpreted as a universal rating of the provider or every deployment of this agent.

Public Findings (1)

Audited security & control findings
CF2Historical Sandbox Escape via Path Traversal in File Patch Tool (Overpatch)
Status: RESOLVED

Flaw in the apply_patch tool permitted directory path manipulation to write outside designated workspace boundaries. Independently reported under OpenAI Bugcrowd CVD program in August 2026, reproduced, and remediated in production release within 8 days prior to assessment snapshot.

Finding ID: CF-CODEX-001 · Control: TR02.02
Critical Control Coverage (30 Controls)
25Applicable
0E3 Tested
6Not Testable
Evidence Distribution (9 Records)
0E0
6E1
3E2
0E3
Operational Profile & Autonomy Boundaries
Autonomy Level:A2
Financial Authority:F0
Guardrails:GD2
Composition Risk:CE1
Cryptographic Assessment Fingerprint:
14612321d34ccd1ed133eb8c5be65d3dc9d028d18270512e13e051d65861b9f7
View Full Trust Passport

Sources & Provenance

(0)
Information observed: 1 October 2026

Registry Provenance Distinction: Registry Source Provenance is NOT TrustRank Assessment Evidence. Sources above document registry facts. TrustRank assessment evidence measures technical control test results.

No public registry provenance is currently available for this profile.