TATECHATLAS
◎ English
Programming / Guide

Serializing Complex Python Objects to JSON with Custom Encoders

The json module handles only a fixed set of Python types. To serialize complex numbers, dates, paths, or user-defined classes you must supply a default function or subclass JSONEncoder, then pair the output with object_hook on the decode side.

On this page

Python's json module converts only dict, list, tuple, str, int, float, bool, and None by default. Anything else raises TypeError unless you intercept it. The interception happens through the default parameter of json.dumps or through the default method of a JSONEncoder subclass. Both approaches receive the unsupported object and must return a JSON-compatible value or raise TypeError to signal failure. For round-trip fidelity you must also provide a matching object_hook on the decode side that reconstructs the original type from the marker dictionary or string.

Why JSON Does Not Serialize Arbitrary Python Objects

The json module defines a fixed conversion table. Python dict maps to a JSON object, list and tuple map to an array, str maps to a string, int and float map to a number, True and False map to their JSON literals, and None maps to null. Complex numbers, datetime instances, pathlib.Path objects, and user-defined class instances have no entry in this table. When the encoder encounters a value outside the table it calls the default hook. If no hook is provided the default hook raises TypeError. The documentation states this explicitly: the encoder supports only the listed types and to extend it you must subclass JSONEncoder and implement a default method.

Source excerpt from the Python documentation: 'To extend this to recognize other objects, subclass and implement a default() method with another method that returns a serializable object for o if possible, otherwise it should call the superclass implementation (to raise TypeError).'

import json
# This raises TypeError because complex has no default mapping
try:
    json.dumps(1 + 2j)
except TypeError as e:
    print(e)

When you serialize a custom class, always embed a type discriminator key such as __type__ so that the decoder can distinguish your marker dictionary from a plain dict that happens to share the same field names.

Serialization via the default Parameter in json.dumps

The default parameter accepts a callable that receives the unsupported object and returns a JSON-serializable value. The return value can be a dict, a list, a string, a number, or a boolean. If the callable cannot handle the object it must raise TypeError so that other objects in the same document can still be encoded by their own rules or by a chained super call. This approach is ideal when you need a one-off transformation in a single call site and do not want to define a reusable class.

The source documentation shows the canonical example with complex numbers: a function checks isinstance(obj, complex), returns a marker dictionary with real and imag fields, and raises TypeError for everything else.

import json

def custom_json(obj):
    if isinstance(obj, complex):
        return {"__complex__": True, "real": obj.real, "imag": obj.imag}
    raise TypeError(f"Cannot serialize object of {type(obj)}")

print(json.dumps(1 + 2j, default=custom_json))

Extending Behavior Through a JSONEncoder Subclass

When the same serialization logic must be reused across multiple call sites or when you need to configure encoder-level parameters such as indent or ensure_ascii together with custom type handling, subclassing JSONEncoder is the cleaner option. You override the default method and call super().default(o) for unhandled types so that the standard TypeError is raised with the library's own message. The subclass is passed to json.dumps or json.dump through the cls parameter.

This pattern separates the encoding policy from the call site. Any consumer that imports your encoder class gets consistent behavior without duplicating the isinstance chain.

import json

class SimpleEncoder(json.JSONEncoder):
    def default(self, o):
        if isinstance(o, complex):
            return {"__complex__": True, "real": o.real, "imag": o.imag}
        return super().default(o)

print(json.dumps({"z": 3 + 4j}, cls=SimpleEncoder))

Choosing Between a default Function and a JSONEncoder Subclass

Use a plain default function when the transformation is specific to one call, the logic is short, and you do not need to share encoder configuration. The function is passed inline and disappears after the call. Use a JSONEncoder subclass when the logic must be reused, when you want to combine custom type handling with non-default encoder settings such as sort_keys or a custom separators tuple, or when the transformation chain grows beyond two or three isinstance checks and a flat function becomes hard to read.

A practical threshold: if you find yourself passing the same default function to dumps in more than two places, promote it to a class. If the class only overrides default and adds no constructor parameters, the function form is still acceptable.

import json
from datetime import datetime, timezone

class ProjectEncoder(json.JSONEncoder):
    def default(self, o):
        if isinstance(o, datetime):
            return o.isoformat()
        if isinstance(o, complex):
            return {"__complex__": True, "real": o.real, "imag": o.imag}
        return super().default(o)

payload = {"ts": datetime(2024, 1, 15, tzinfo=timezone.utc), "z": 1+1j}
print(json.dumps(payload, cls=ProjectEncoder, sort_keys=True))

Encoding Dates, Paths, and User-Defined Classes

datetime objects have no JSON mapping. The conventional representation is the ISO 8601 string produced by the isoformat method. The decoder side uses datetime.fromisoformat to reconstruct the value. pathlib.Path objects serialize to their string form via str(path); the decoder wraps the string back with Path(). For user-defined classes the recommended pattern is a marker dictionary that includes a type discriminator key, for example __type__ set to the class name, plus the fields needed for reconstruction.

Consistency between encode and decode is critical. If the encoder emits {"__type__": "Point", "x": 1, "y": 2} the decoder's object_hook must check for __type__ equal to "Point" and return the corresponding instance. Without a discriminator key a plain dict with the same shape would be indistinguishable from your custom object.

import json
from datetime import datetime, timezone
from pathlib import Path

class Point:
    def __init__(self, x, y):
        self.x = x
        self.y = y

def encode(obj):
    if isinstance(obj, datetime):
        return obj.isoformat()
    if isinstance(obj, Path):
        return str(obj)
    if isinstance(obj, Point):
        return {"__type__": "Point", "x": obj.x, "y": obj.y}
    raise TypeError(f"Unsupported type {type(obj)}")

def decode(dct):
    if dct.get("__type__") == "Point":
        return Point(dct["x"], dct["y"])
    return dct

p = Point(3, 4)
s = json.dumps(p, default=encode)
print(s)
restored = json.loads(s, object_hook=decode)
print(type(restored).__name__, restored.x, restored.y)

Encoder Parameters That Affect Output

ensure_ascii controls whether non-ASCII characters are escaped as \uXXXX sequences or emitted as-is. When your custom default returns strings containing non-ASCII text this setting changes the byte representation but not the logical content. indent adds whitespace for readability and changes the default separators from (',', ':') to (', ', ': '). sort_keys sorts dictionary keys alphabetically before string conversion, which is useful for deterministic output in tests. separators lets you produce compact output by passing (',',':'). allow_nan governs whether NaN, Infinity, and -Infinity are encoded as JavaScript literals or raise ValueError. check_circular detects reference cycles in containers and raises ValueError when enabled; if disabled a cycle causes RecursionError. skipkeys silently drops dictionary keys that are not str, int, float, bool, or None instead of raising TypeError.

Source excerpt: 'If check_circular is true (the default), then lists, dicts, and custom encoded objects will be checked for circular references during encoding to prevent an infinite recursion.'

import json
from datetime import datetime

class Enc(json.JSONEncoder):
    def default(self, o):
        if isinstance(o, datetime):
            return o.isoformat()
        return super().default(o)

print(json.dumps({"b": 1, "a": datetime(2024,1,1)}, cls=Enc, sort_keys=True, indent=2))

Error Handling and Portability Limitations

When default or the encoder's default method raises TypeError the entire dumps call fails; there is no partial output. This means one unsupported object anywhere in the tree aborts the whole serialization. Design your default function to raise TypeError with a clear message naming the offending type so that debugging is fast.

JSON is not a framed protocol. Repeated calls to json.dump with the same file object produce concatenated documents that are not valid JSON. Keys that are not strings become strings after round-trip, so loads(dumps(x)) may not equal x if x had integer or tuple keys. Very large integers and decimal.Decimal values can exceed the precision of IEEE 754 double-precision consumers. The documentation warns that malicious JSON can cause the decoder to consume considerable CPU and memory, so limit input size when parsing untrusted data.

Source excerpt: 'Unlike pickle and marshal, JSON is not a framed protocol, so trying to serialize multiple objects with repeated calls to dump() using the same fp will result in an invalid JSON file.'

import json, io
buf = io.StringIO()
json.dump({"a": 1}, buf)
json.dump({"b": 2}, buf)  # second call makes the file invalid
print(buf.getvalue())  # {"a": 1}{"b": 2} - not valid JSON

Practical Example Combining Complex Numbers and a Custom Class

The following complete example demonstrates both approaches side by side. A standalone default function handles complex numbers at one call site. A JSONEncoder subclass handles complex numbers and a custom Event class at another call site, combined with sort_keys for deterministic output. The decode side uses object_hook to reconstruct both types from their marker dictionaries.

This example runs on Python 3.6 and later. It requires no third-party packages. The only prerequisite is familiarity with isinstance, dictionaries, and the json module's basic API.

import json

class Event:
    def __init__(self, name, ts, payload):
        self.name = name
        self.ts = ts
        self.payload = payload

# Approach 1: standalone default function
def complex_default(obj):
    if isinstance(obj, complex):
        return {"__complex__": True, "real": obj.real, "imag": obj.imag}
    raise TypeError(f"Cannot serialize {type(obj)}")

s1 = json.dumps(1 + 2j, default=complex_default)
print(s1)

# Approach 2: reusable JSONEncoder subclass
class AppEncoder(json.JSONEncoder):
    def default(self, o):
        if isinstance(o, complex):
            return {"__complex__": True, "real": o.real, "imag": o.imag}
        if isinstance(o, Event):
            return {"__type__": "Event", "name": o.name,
                    "ts": o.ts, "payload": o.payload}
        return super().default(o)

s2 = json.dumps({"z": 3+4j, "ev": Event("click", 100, {"x": 1})},
                cls=AppEncoder, sort_keys=True)
print(s2)

# Round-trip decode
def app_hook(dct):
    if dct.get("__complex__"):
        return complex(dct["real"], dct["imag"])
    if dct.get("__type__") == "Event":
        return Event(dct["name"], dct["ts"], dct["payload"])
    return dct

restored = json.loads(s2, object_hook=app_hook)
print(type(restored["z"]).__name__, restored["ev"].name)

Things to check

  • json.dumps with no default or cls raises TypeError for complex, datetime, Path, or arbitrary class instances
  • A default function must raise TypeError for unhandled types; returning None silently encodes null
  • A JSONEncoder subclass must call super().default(o) for unhandled types to preserve the standard error message
  • Round-trip fidelity requires a matching object_hook on the decode side that checks the marker key
  • Repeated json.dump calls to the same file object produce invalid concatenated output
  • ensure_ascii, indent, sort_keys, and separators affect representation but not the logical content produced by default
  • allow_nan=False turns NaN and Infinity encoding into ValueError
  • check_circular=True detects cycles in custom-encoded objects and raises ValueError

This article covers only the standard library json module on CPython 3.6 and later. It does not address third-party serializers such as orjson, ujson, or rapidjson which have their own extension mechanisms. The round-trip guarantees discussed here depend entirely on the consistency between your encode markers and decode hooks; the library itself provides no type registry or schema validation. Very large integers and decimal.Decimal values may lose precision in JSON consumers that parse numbers as IEEE 754 doubles. Input size limits for untrusted JSON are the application's responsibility, not the module's.

Sources

  1. Python: json ↗
  2. Python: pathlib ↗
  3. Python: os and working directories ↗
Back to top ↑