Working with HTTP Resources

Overview

OCM can add a file served over HTTP or HTTPS to a component version using the Wget/v1 type. There is one download engine underneath, with two front doors:

  • The input type downloads the file while the component version is built and stores the bytes inside it.
  • The access type stores only the URL and leaves the bytes on the remote server.

They share the same download, credential, and configuration code, so the choice between them is about where the bytes live, not about what you can configure.

If you just want the commands to add a file and download it back, start with the How-To guide Add Resources from HTTP URLs. This tutorial explains the parts that guide leaves out: why you’d pick one type over the other, how credentials are matched, how the media type is decided, how digest pinning protects you, and what changed since OCM v1.

What You’ll Learn

  • When to embed a file (input type) and when to reference it by URL (access type)
  • What happens to a Wget resource when you transfer the component version
  • How OCM resolves a resource’s media type
  • How credentials are matched to a request, and how the three authentication methods interact
  • How to pin a resource by digest so a changed file is rejected instead of silently accepted
  • How to send a non-GET request and tune timeouts, retries, and the download directory
  • What changed between OCM v1 and OCM v2, and how to migrate

Prerequisites

  • OCM CLI installed
  • Comfortable adding a component version with ocm add cv; the how-to is a good warm-up

Input or access: where the bytes live

Both types describe the same resource with the same fields. The only difference is when the file is fetched and where it ends up.

  • With the input type, OCM downloads the file the moment you run ocm add cv and stores it inside the component version as a local blob. The original URL is not kept. The component version is now self-contained.

  • With the access type, OCM stores the URL in the component descriptor and fetches the bytes only when someone asks for them. The URL stays the source of truth, which is what you want for a large file, or when consumers are meant to pull directly from the origin server.

The rule of thumb is: use input type when you want a reproducible copy, and the access type when the URL is the authoritative location and should be used for accessing the resource.

How transfer works

By default, a Wget resource stays by reference: without a matching uploader, ocm transfer cv keeps the Wget/v1 access unchanged in the target, and the file stays on the remote server.

With a matching local blob uploader configuration (in your OCM configuration, for example .ocmconfig in the working directory), OCM fetches the bytes and writes them into the target as a LocalBlob/v1. This means:

  1. To embed Wget resources, add a localblob.uploader.transfer.config.ocm.software/v1alpha1 entry to your OCM configuration. Without it, the Wget/v1 access remains by reference.
  2. When the local blob uploader copies the resource, the bytes are fetched again at transfer time and checked against the resource’s digest. If the file behind the URL changed since the component version was built, the transfer fails instead of copying different content.

To keep a resource behind a URL instead of embedding it — streaming it to a custom HTTP target and rewriting the access to a new Wget/v1 URL — configure an uploader; see Configure Custom Uploads During Transfer.

Forward a digest header to the upload target

An uploader can template its header values as ${…} CEL expressions over the source resource, so the upload PUT can carry a checksum the source already advertised, read from resource.digest. To emit a standards-compliant Content-Digest (RFC 9530), map the OCM algorithm name to the RFC 9530 key with contentDigestAlgorithm() and convert the hex digest to base64 with base64.encode(hex.decode(...)):

type: generic.config.ocm.software/v1
configurations:
  - type: http.uploader.transfer.config.ocm.software/v1alpha1
    match: resource.access.isType("Wget/v1")
    targetURL: '${"https://mytarget.example.com/uploads" + url(resource.access.url).path}'
    method: PUT
    header:
      # RFC 9530 Content-Digest: sha-256=:<base64>:
      Content-Digest: ['${contentDigestAlgorithm(resource.digest.hashAlgorithm) + "=:" + base64.encode(hex.decode(resource.digest.value)) + ":"}']

Use the Content-Digest expression only when resource.digest was produced by a byte-preserving normalization such as genericBlobDigest/v1. The uploader streams the source bytes unchanged, so a digest built from a non-byte-preserving normalization does not describe the transmitted content and a target that validates Content-Digest may reject the upload.

resource.digest is only present when the source resource carries a digest (for example when it is pinned from the source via the checksum-http configuration). See Templating Headers for the full field reference.

For a target such as JFrog Artifactory, which verifies uploads against a hex X-Checksum-* header, resource.digest.value maps on directly — no base64 conversion:

  - type: http.uploader.transfer.config.ocm.software/v1alpha1
    match: resource.access.isType("Wget/v1")
    targetURL: '${"https://myorg.jfrog.io/artifactory/my-repo" + url(resource.access.url).path}'
    method: PUT
    header:
      X-Checksum-Sha256: ['${resource.digest.value}']

Set the media type

For a Wget resource, OCM picks the media type in the following order and stops at the first one it finds:

  1. The mediaType field in your specification, if you set it.
  2. The Content-Type header the server returns.
  3. application/octet-stream, as a last resort.

Practical advice: set mediaType yourself whenever the server doesn’t return a useful Content-Type. The most common examples are GitHub release assets. GitHub serves all of them as application/octet-stream, no matter what they actually are, so a .tar file would end up with a meaningless media type unless you set mediaType: application/x-tar explicitly.

OCM v2 does not guess the media type from the file extension in the URL. OCM v1 did. See Migrate from OCM v1 if you are moving old constructor files across.

Authentication

Credentials are resolved through OCM’s normal credential system. Never add authentication to a URL directly. OCM builds a consumer identity of type Wget from the URL and uses the matching consumer entry’s credentials for the request. Because the input type and the access type build that identity the same way, one entry will work for both, building the component version, and downloading the resource. For how identities are matched and resolved, see Understand Credential Resolution.

Wget’s credential type is: WgetCredentials/v1.

Authentication methods

WgetCredentials/v1 supports three methods:

MethodFieldsSets
HTTP Basic Authusername, passwordAuthorization: Basic …
Bearer tokenidentityTokenAuthorization: Bearer …
Mutual TLS (client cert)certificate, privateKey, and optional certificateAuthorityThe TLS handshake
  • Basic Auth and a bearer token both set the Authorization header, so they are mutually exclusive. If you configure both, the bearer token wins and OCM logs a warning.
  • The client certificate is separate. It is applied during the TLS handshake, not in a header, so it works independent of the other two.

Put secrets in your credential configuration, never in the specification. Anything you write into url, header, or body (including https://user:token@host/... and presigned query parameters) is stored with the component version (access type) or lives in your constructor file (input type). Neither is a safe place for a secret.

For the full field reference, see Credential Types: WgetCredentials/v1.

Send a non-GET request

By default, OCM sends a GET request. To send something else, set verb, and optionally header and body. body is base64-encoded in YAML, because the underlying field is a byte slice:

resources:
  - name: report
    type: blob
    version: 1.0.0
    input:
      type: Wget/v1
      url: https://api.example.com/reports
      verb: POST
      mediaType: application/json
      header:
        Accept:
          - application/json
        X-Request-Source:
          - ocm
      # base64 of {"format":"json"}
      body: eyJmb3JtYXQiOiJqc29uIn0=

Fine-tuning the download

Only successful responses are accepted. OCM accepts a 2xx status and fails on anything else. Redirects are followed by default. If you set noRedirect: true, the request fails.

Downloads are streamed to disk, not held in memory. The response body is written straight to a temporary file, so memory use stays flat no matter how big the file is. There is no size limit by default. That temporary file is created under the tempFolder of the filesystem.config.ocm.software/v1alpha1 attribute in the .ocmconfig configuration file, falling back to the operating system’s temp directory if no configuration is provided.

Timeouts, retries, and per-host settings come from the http.config.ocm.software/v1alpha1 configuration and apply to every Wget request:

type: generic.config.ocm.software/v1
configurations:
  - type: filesystem.config.ocm.software/v1alpha1
    tempFolder: /var/tmp/ocm
  - type: http.config.ocm.software/v1alpha1
    timeout: 2m
    retry:
      maxRetries: 3
    hosts:
      github.com:
        timeout: 5m

See HTTP Client Configuration for the full schema, defaults, and how per-host settings are merged.

Verifying downloads against the source

A pinned resource digest protects consumers after the build, but the first ocm add cv still trusts whatever the server sends. A source-side checksum closes that gap: OCM compares the downloaded bytes against a digest the source advertises in response headers (RFC 9530 Content-Digest, x-checksum-*) and fails on a mismatch. It is configured centrally via checksum.http.config.ocm.software/v1alpha1, not in the constructor. See HTTP Checksum Configuration for the schema, checksum modes, precedence, and the access-side fast path.

Migrate from OCM v1

Three things changed between OCM v1 and v2: credential matching, constructor syntax, and behavior. The credential changes are the most likely to break existing configurations.

Credential changes

Identity matching has been updated.

FieldOCM v1OCM v2What to do
Consumer identitytype: wgettype: WgetRename the type.
Identity pathpathprefix, longest-prefix matchpath, glob matchRename the attribute. * matches one path segment; omit path to match the whole host.

Further matching changes:

AreaOCM v1OCM v2
Auth precedenceBasic Auth wins; the bearer token is a fallbackBearer token wins; Basic Auth is used only when no token is set
Basic AuthNeeds both username and passwordusername is enough
Custom CARoot CAs from credentials always appliedcertificateAuthority applied only alongside certificate

Constructor changes

FieldOCM v1OCM v2What to do
input.bodyPlain stringBase64-encoded byte sliceBase64-encode the body in the constructor.

In OCM v1 the body was an io.Reader, which has no YAML form. In OCM v2 it is a byte slice, therefore it needs to be base64.

Behavior changes

AreaOCM v1OCM v2
Media typemediaType → Content-Type → URL file extension → application/octet-streammediaType → Content-Type → application/octet-stream
Minimum TLSTLS 1.3TLS 1.2

OCM v2 dropped the file-extension guess, so a .tar.gz URL that used to resolve to application/x-gzip on its own, now needs an explicit mediaType.

Troubleshooting

401 Unauthorized from the server

Why: Either the credentials are wrong, or no consumer entry matched the identity OCM built from the URL. The second case is hard to spot: when nothing matches, OCM sends the request without credentials and the server returns the same 401 either way.

Fix: Check the credentials first, then check that the entry matches. hostname must equal the URL host; scheme, port, and path narrow the match only when set, so the broadest entry is the one that sets only hostname. See Credential Consumer Identities: Wget.

401 Unauthorized right after migrating from OCM v1

Why: Two identity attributes were renamed. type: wget MUST become type: Wget, and pathprefix no longer exists, because OCM v2 uses path.

Fix: Rename the type and drop pathprefix. Omitting path matches every path on the host. See Credential changes.

The resource has media type application/octet-stream

Why: No mediaType was set and the server sent no useful Content-Type. OCM v2 does not fall back to the file extension.

Fix: Set mediaType on the input or access specification.

digest mismatch when adding the component version

Why: The resource pins a digest and the fetched bytes hash to something else: the content behind the URL changed, the download was truncated, or the pinned value came from a different file. The error to watch for is resource blob digest mismatch.

Fix: Re-download the URL and recompute the digest. If the new value is the one you have, update value. If it isn’t, the content changed.

The operation fails reporting a 3xx status

Why: noRedirect: true is set.

Fix: Remove noRedirect, or point url at the final location the redirect resolves to.