Standard Codecs and Shared Vocabulary
The stock conversions for the standard families, and the mix-in pattern that shares them across an API.
A typical DTO boundary converts the same handful of families every time: identifiers, dates, enums, money. Writing a ValidatedPrism by hand for each would be busywork, and writing it lawfully, accepting exactly the spelling it renders, is subtle. StandardCodecs ships that vocabulary ready-made. This page first walks the stock codecs and the one rule they all keep. Your own canon and a vocabulary shared across specs follow, for when you need them.
- Map identifiers, dates, enums, numbers and money with one factory call each, and predict which spellings each accepts
- Predict which of a browser's or Python's timestamps a stock codec rejects, and declare the producer's canon instead
The code on this page is StandardCodecsBook.java and its StandardCodecsBookTest.java - the page includes them directly, so they are compiled and run by the build.
Standard codecs
The common conversion families need no conversion code of your own. StandardCodecs, in org.higherkindedj.optics.validated, ships one factory per family. The processor never applies one on its own, so each converting component declares a one-line leaf that returns it:
import static org.higherkindedj.optics.validated.StandardCodecs.bigDecimal;
import static org.higherkindedj.optics.validated.StandardCodecs.enumByName;
import static org.higherkindedj.optics.validated.StandardCodecs.instant;
import static org.higherkindedj.optics.validated.StandardCodecs.localDate;
import static org.higherkindedj.optics.validated.StandardCodecs.offsetDateTime;
import static org.higherkindedj.optics.validated.StandardCodecs.uuid;
enum OrderStatus {
NEW,
PAID,
SHIPPED
}
record Order(UUID id, LocalDate placedOn, OrderStatus status, BigDecimal total) {}
record OrderDto(String id, String placedOn, String status, String total) {}
@GenerateMapping
interface OrderMapping extends MappingSpec<Order, OrderDto> {
default ValidatedPrism<String, UUID> id() {
return uuid();
}
default ValidatedPrism<String, LocalDate> placedOn() {
return localDate();
}
default ValidatedPrism<String, OrderStatus> status() {
return enumByName(OrderStatus.class);
}
default ValidatedPrism<String, BigDecimal> total() {
return bigDecimal();
}
}
| Factory | Wire ↔ domain |
|---|---|
uuid() | String ↔ UUID |
uri() | String ↔ URI |
localDate() / localDate(DateTimeFormatter) | String ↔ LocalDate |
instant() | String ↔ Instant (UTC, Z) |
offsetDateTime() / offsetDateTime(DateTimeFormatter) | String ↔ OffsetDateTime |
enumByName(Class) | String ↔ any enum, by exact constant name |
bigDecimal() | String ↔ BigDecimal, plain notation, scale preserved |
intFromString() / longFromString() / doubleFromString() | String ↔ boxed number, canonical spellings only ("2" is not a canonical double; "2.0" is) |
booleanStrict() | String ↔ Boolean, exactly true/false |
currency() | String ↔ Currency (ISO 4217) |
locale() | String ↔ Locale (BCP 47 tag) |
A codec's own failure is a FieldError with a message and no path. Under a spec, the generated parse locates it at the component, so the codecs feed the 422 leg unchanged. Each message ends with a sample of the spelling the codec wants, and the enum message names the permitted constants:
Validated<NonEmptyList<FieldError>, Order> parsed =
OrderMappingImpl.INSTANCE.parse(new OrderDto("NOPE", "28/07/2026", "DISPATCHED", "1E+3"));
assertThatValidated(parsed)
.isInvalid()
.hasFieldErrors(
"id: not a UUID (expected e.g. 123e4567-e89b-12d3-a456-426614174000)",
"placedOn: not an ISO-8601 date (expected e.g. 2026-07-28)",
"status: unknown OrderStatus (expected one of NEW, PAID, SHIPPED)",
"total: not a number in plain notation (expected e.g. 123.45)");
The number and boolean codecs are for values that arrive as strings: a query parameter, a CSV cell, a quoted JSON value. They produce boxed types, so the domain component is Integer or Boolean, not int: a ValidatedPrism cannot name a primitive, and the processor refuses the mismatch. A quantity that JSON sends as a number can take either of two shapes. Declared Integer on both sides, it needs no codec, and a missing one is still a located must not be null. But Jackson answers a malformed one before parse runs, with its own 400 for "two", and it quietly truncates 2.5 to 2. Declared String on the wire with intFromString(), as in the Capstone, every bad quantity is located, and a JSON number 2 still binds, as the string "2".
A leaf for a component called currency, locale, uri or uuid shares its factory's name. Inside default ValidatedPrism<String, Currency> currency(), an unqualified currency() calls the leaf itself, not the statically imported factory, so the first parse or build throws a StackOverflowError. It compiles without a warning. Write return StandardCodecs.currency();.
Canonical forms only
A codec's canon is the spelling its build writes. parse accepts that spelling and rejects every other spelling of the same value with an error, honouring the section law: an accepted wire value must rebuild to exactly itself. So parse rejects an uppercase UUID, 042, 1E+3 or the language tag en-gb, and never quietly normalises it. The rejection reads like a malformed value, so its expected e.g. sample is what shows the client the spelling the codec wants.
Silent normalisation is data mutation nobody asked for. A mapper that quietly lowercases a UUID or reformats a timestamp makes an echo endpoint return different bytes than it received. It breaks cache keys and payload signatures, and bakes a client's spelling bug into the contract without anyone deciding to. Strictness is what guarantees a round trip: every stock codec is tested to accept exactly what it writes. When a producer speaks a different canon, you do not weaken the rule; you declare that canon, and keep the same guarantee on their spelling.
The two date-time codecs render differently, so they accept different spellings of the same moment. Both write UTC as Z, and leave out a fraction that is zero. instant() writes as Instant.toString() does: always UTC, and any other fraction in groups of three digits. offsetDateTime() keeps the offset it was given, and drops the fraction's trailing zeros:
| Spelling on the wire | Accepted by instant() | Accepted by offsetDateTime() |
|---|---|---|
2026-07-28T12:34:56Z | ✅ | ✅ |
2026-07-28T12:34:56.500Z | ✅ | ❌ |
2026-07-28T12:34:56.5Z | ❌ | ✅ |
2026-07-28T12:34:56.123Z | ✅ | ✅ |
2026-07-28T12:34:56.000Z | ❌ | ❌ |
2026-07-28T12:34:56+00:00 | ❌ | ❌ |
2026-07-28T12:34:56+01:00 | ❌ | ✅ |
A browser's toISOString() always writes three fraction digits and Z. instant() accepts that except when the milliseconds are zero, since Instant.toString() never writes .000Z. For a timestamp read from the clock, that is about one request in a thousand. For a time the user picked in a date or time input, it is every request, because those are whole seconds. An aware Python datetime's isoformat() writes +00:00 for UTC, which both codecs reject every time. Nothing fails at compile time, and a test whose fixture has non-zero milliseconds passes every time. Declare the producer's canon instead, and check what your producer really sends: some Python frameworks rewrite +00:00 as Z.
For a LocalDate or OffsetDateTime component, the formatter overload is enough: pass the producer's pattern, and the codec accepts exactly what that pattern writes. instant() has no formatter overload, and no single pattern describes Python's output. For those, ValidatedPrism.canonical wraps a parser and a render, and rejects every spelling the render would not write back identically:
final class WireFormats {
// A browser's toISOString(): always three fraction digits, and Z. withZone lets it write an
// Instant.
private static final DateTimeFormatter BROWSER =
DateTimeFormatter.ofPattern("uuuu-MM-dd'T'HH:mm:ss.SSSX").withZone(ZoneOffset.UTC);
// For an Instant component. instant() rejects .000Z, which Instant.toString() never writes.
static final ValidatedPrism<String, Instant> BROWSER_INSTANT =
ValidatedPrism.canonical(
"not a browser timestamp (expected e.g. 2026-07-28T12:34:56.000Z)",
raw -> Instant.from(BROWSER.parse(raw)),
BROWSER::format);
// For an OffsetDateTime component, the formatter overload takes the browser's spelling, at any
// offset.
static final ValidatedPrism<String, OffsetDateTime> BROWSER_OFFSET =
offsetDateTime(DateTimeFormatter.ofPattern("uuuu-MM-dd'T'HH:mm:ss.SSSXXX"));
// Python's isoformat(): +00:00 for UTC, and six fraction digits, or none when they are zero.
private static final DateTimeFormatter WHOLE_SECONDS =
DateTimeFormatter.ofPattern("uuuu-MM-dd'T'HH:mm:ssxxx");
private static final DateTimeFormatter MICROSECONDS =
DateTimeFormatter.ofPattern("uuuu-MM-dd'T'HH:mm:ss.SSSSSSxxx");
static final ValidatedPrism<String, OffsetDateTime> PYTHON_OFFSET =
ValidatedPrism.canonical(
"not a Python isoformat() timestamp (expected e.g. 2026-07-28T12:34:56.123456+00:00)",
OffsetDateTime::parse,
time -> time.format(time.getNano() / 1_000 == 0 ? WHOLE_SECONDS : MICROSECONDS));
// For an Instant component: the same two shapes, with the offset fixed at +00:00.
static final ValidatedPrism<String, Instant> PYTHON_INSTANT =
ValidatedPrism.canonical(
"not a Python isoformat() UTC timestamp (expected e.g. 2026-07-28T12:34:56.123456+00:00)",
raw -> OffsetDateTime.parse(raw).toInstant(),
instant -> PYTHON_OFFSET.build(instant.atOffset(ZoneOffset.UTC)));
private WireFormats() {}
}
Pick the leaf by producer and component type, and name it after the component, as usual:
| Producer | Instant component | OffsetDateTime component |
|---|---|---|
a browser's toISOString() | BROWSER_INSTANT | BROWSER_OFFSET |
Python's isoformat(), on an aware datetime | PYTHON_INSTANT | PYTHON_OFFSET |
default ValidatedPrism<String, Instant> placedAt() { return WireFormats.BROWSER_INSTANT; } accepts every spelling a browser sends, and renders a parsed value back the same way. One leaf has one canon, since build writes one spelling. When a browser and a Python job send the same record, declare a spec per producer over the same pair, each with its own timestamp leaf, and share the rest through a vocabulary.
BROWSER_INSTANT writes milliseconds, and PYTHON_INSTANT microseconds. An Instant.now() carries finer digits on current JDKs, so build cuts them off, and a client that echoes the value back sends a different Instant. Truncate at creation where a value must round-trip: Instant.now().truncatedTo(ChronoUnit.MILLIS). A law check from a wire sample never sees the finer value, but ValidatedPrismLaws.assertParseBuild(WireFormats.BROWSER_INSTANT, Instant.now()) does, and fails.
You can now map the standard families with one factory call each, predict which spellings a codec accepts, and take a browser's or Python's timestamps without loosening anything. The rest of this page, your own canon and a vocabulary shared across specs, is for when you need them.
A quantity component is mapped with StandardCodecs.intFromString(). Clients send "42", "042", "+42", " 42" and "42.0". Which of them parse?
- All five: each names the number 42
"42","042"and"+42", the onesInteger.parseIntaccepts"42"only"42"and"42.0"
Answer and why
Answer and why
3. A codec accepts exactly the spelling it renders, and Integer.toString writes 42 and nothing else. Integer.parseInt would take "042" and "+42", but build could never reproduce either, so parse rejects both with a located error:
ValidatedPrism<String, Integer> quantity = StandardCodecs.intFromString();
assertThatValidated(quantity.parse("42")).isValid();
for (String spelling : List.of("042", "+42", " 42", "42.0")) {
assertThatValidated(quantity.parse(spelling)).isInvalid(); // Integer.parseInt takes two
}
Where this lives: Canonical forms only.
A browser stamps each request with new Date().toISOString(), and an OffsetDateTime component maps it with the plain StandardCodecs.offsetDateTime(). Roughly how often does parse reject a request?
- Never: both write UTC as
Z - About one request in a thousand, as for
instant() - About one request in ten
- Every time:
offsetDateTime()expects an offset such as+01:00
Answer and why
Answer and why
3. offsetDateTime() drops the fraction's trailing zeros when it renders, so it writes .12Z where a browser sends .120Z. It rejects every millisecond value that ends in a zero, a hundred of the thousand, while BROWSER_OFFSET takes them all:
List<String> sent = // every millisecond value, as toISOString() writes it
IntStream.range(0, 1000).mapToObj("2026-07-28T12:34:56.%03dZ"::formatted).toList();
List<String> rejected =
sent.stream().filter(spelling -> offsetDateTime().parse(spelling).isInvalid()).toList();
assertThat(rejected).hasSize(100).allMatch(spelling -> spelling.endsWith("0Z"));
assertThat(sent).allMatch(spelling -> WireFormats.BROWSER_OFFSET.parse(spelling).isValid());
Where this lives: Canonical forms only.
Your own canon
Declaring the producer's canon, as the browser and Python leaves do, covers any wire with a canon of its own. An uppercase-UUID producer (SQL Server) is not forbidden by the law; only accepting both cases through one leaf is. The lenient, throwing UUID.fromString is fine inside ValidatedPrism.canonical, because the render defines the canon:
// A pallet from the warehouse's feed, which writes UUIDs in upper case, as SQL Server does.
record Pallet(UUID id, String label) {}
record PalletDto(String id, String label) {}
@GenerateMapping
interface PalletMapping extends MappingSpec<Pallet, PalletDto> {
// The lenient, throwing parse is fine: the render defines the canon, and the per-value guard
// rejects every spelling it cannot reproduce.
default ValidatedPrism<String, UUID> id() {
return ValidatedPrism.canonical(
"not an uppercase UUID",
UUID::fromString,
uuid -> uuid.toString().toUpperCase(Locale.ROOT));
}
}
A pattern's canon is everything the pattern writes, not just what one producer sends. So BROWSER_OFFSET also accepts a +01:00 offset a browser never sends, lawfully, and rejects extra fraction digits, which the pattern has no room for.
Conversions the vocabulary does not cover stay hand-written leaves: ValidatedPrism.canonical(...) where a throwing parser and a render exist, ValidatedPrism.of(...) for full control.
A parse given to canonical may also answer null, which is a rejection that costs no exception. So a leaf that meets malformed input often can check the source's shape first, and answer null for a bad one, before a JDK parser builds a stack trace. The stock codecs do this.
The generated Impl reads each leaf once, on first use, and keeps what it answers for every caller, on every thread. So a leaf that picks its codec on each call, from a flag, a system property or a field a test changes, keeps its first pick. Make a choice that must vary inside the codec's parse. A codec built over SimpleDateFormat or DecimalFormat would be shared by threads that each assume it is theirs, so build over DateTimeFormatter instead.
// enumByName builds a new codec on every call...
assertThat(StandardCodecs.enumByName(OrderStatus.class))
.isNotSameAs(StandardCodecs.enumByName(OrderStatus.class));
// ...but the Impl answers the one its status() leaf built first, every time
assertThat(OrderMappingImpl.INSTANCE.status()).isSameAs(OrderMappingImpl.INSTANCE.status());
Shared vocabulary: mix-in interfaces
The same rename or leaf tends to recur across an API's specs: every wire calls it fullName, and every email parses the same way. Record Mapping Basics' PartnerCustomerMapping declared that rename and the email leaf by hand, and every other spec whose wire spells them the same way would repeat both. Move the shared members onto a plain interface, and extend it alongside MappingSpec. Like a Jackson mix-in, it holds mapping declarations apart from the types they describe. Unlike one, a spec extends it, so Java's inheritance decides which declaration wins:
// Plain vocabulary - not a spec itself. Any spec whose records share these
// shapes extends it alongside MappingSpec.
interface ContactVocabulary {
@MapField(to = "fullName")
String name();
default ValidatedPrism<String, EmailAddress> email() {
return EmailCodecs.EMAIL;
}
// Customer has no phone, so this leaf stays inert there; only Supplier binds it.
default ValidatedPrism<String, String> phone() {
return ValidatedPrism.of(
raw ->
raw.matches("\\+?[0-9 ]+")
? Validated.validNel(raw)
: Validated.invalidNel(FieldError.of("not a phone number")),
phone -> phone);
}
}
// A CRM's contact export: the Customer from Record Mapping Basics, calling its name fullName.
record CustomerContactDto(String fullName, String email) {}
@GenerateMapping
interface CustomerContactMapping
extends ContactVocabulary, MappingSpec<Customer, CustomerContactDto> {}
record Supplier(String name, EmailAddress email, String phone) {}
record SupplierDto(String fullName, String email, String phone) {}
@GenerateMapping
interface SupplierMapping extends ContactVocabulary, MappingSpec<Supplier, SupplierDto> {}
// One vocabulary, two mappings: the rename and email leaf apply to both, phone to Supplier.
Validated<NonEmptyList<FieldError>, Customer> customer =
CustomerContactMappingImpl.INSTANCE.parse(
new CustomerContactDto("Ada Lovelace", "not-an-email"));
// Invalid(NonEmptyList[email: not an email address])
Validated<NonEmptyList<FieldError>, Supplier> supplier =
SupplierMappingImpl.INSTANCE.parse(
new SupplierDto("Acme Ltd", "sales@acme.example", "call us"));
// Invalid(NonEmptyList[phone: not a phone number])
An inherited member binds exactly as if it were declared on the spec. One that binds to nothing, like phone() on CustomerContactMapping, stays inert, so one vocabulary serves specs whose records differ. The processor refuses the same member declared on the spec itself, which is what catches a typo there. The flip side: a misspelt leaf in the mix-in stays inert everywhere, so test each spec's rejections once. What an inherited member binds against says what "nothing" means for each kind of member. How a spec collects its vocabulary covers precedence, generic specs and PATCH specs, and Mix-in shapes the processor refuses names the two it will not take.
A vocabulary also crosses a module boundary. The API module that publishes it needs the HKJ libraries its members use, not the annotation processor, and a service module's spec extends it from the jar as if it were local. One API module can therefore own the house vocabulary every service's specs extend, as long as it is on each consumer's compile classpath (Multi-module builds).
- The standard families are one factory call each:
StandardCodecscovers identifiers, dates, enums, numbers and money with lawful codecs, whose failures a spec locates at the component - Canonical forms only: each codec accepts exactly the spelling it renders, so a producer with its own canon gets a leaf that declares it
- Browser and Python timestamps need their own leaves: the stock date-time codecs reject some or all of what they send, and a pattern cuts what it builds to its precision
- Mix-ins share the vocabulary: one plain interface serves every spec, and a member a spec cannot use stays inert
- Validated Prisms: The section law the codecs are built to honour
- Record Mapping Basics: How leaves attach to a spec
- Sparse PATCH: The PATCH sibling that lifts the same element leaves
Previous: Record Mapping Basics Next: Absent Fields and Record Invariants