Cutting JSON Parsing CPU Cost in Go
JSON parsing was a real chunk of our request-handling CPU time. We benchmarked encoding/json, jsoniter, easyjson, sonic, and Go 1.27's encoding/json/v2 against the same payload - decode and encode, on two different machines - instead of trusting each library's own claims.
Full benchmark code: gist.github.com/SergeiSkv/39eeb39d89ee8cef10ed7da6b5aed35a
The JSON CPU Tax
Our API was handling 50M requests/day. JSON parsing showed up as a real, measurable chunk of CPU time in production profiles - a good enough reason to actually measure the alternatives instead of picking one on reputation.
encoding/json relies heavily on reflection and, for a struct with a nested slice, a map, and a time.Time field like the one below, that typically means more allocations than a codegen-based or JIT-based decoder. That's a workload-specific statement, not a universal one - for a flat struct of a few scalar fields, the gap narrows a lot.
The Benchmark
Environment. Go 1.27.1, default build (GOEXPERIMENT unset - see the note on encoding/json/v2 below for why that matters), go test -bench=. -benchmem -count=5, on two machines: an Apple M5 (darwin/arm64) and a Hetzner VPS with an AMD EPYC-Rome vCPU (linux/amd64). Payload is a ~2KB Order struct - 12 line items, a map[string]interface{} metadata field, a time.Time - matching the shape below. Library versions: json-iterator/go v1.1.12, bytedance/sonic v1.15.3, mailru/easyjson v0.9.2, tidwall/gjson v1.19.0. Full benchmark code is on GitHub Gist, so you can run it against your own payload instead of trusting these numbers on faith. On reproducibility: repeated benchstat runs put per-library run-to-run spread at roughly ±1-2% on both machines. A gap below that - like sonic's M5 encode result (3.11µs vs 3.10µs) or jsoniter's M5 decode result (4.35µs vs 4.29µs) - is noise, not a finding, and is called out as such below. A gap clearly past that band - jsoniter's ~7% slower VPS decode, or anything in the 8%+ range elsewhere in this article - held up across multiple re-runs and is reported as real.
type Order struct {
ID string `json:"id"`
UserID int64 `json:"user_id"`
Items []OrderItem `json:"items"`
Total float64 `json:"total"`
CreatedAt time.Time `json:"created_at"`
Metadata map[string]interface{} `json:"metadata"`
}
Decode (unmarshal bytes into the struct - the direction "JSON parsing" usually means):
| Library | M5 | VPS (EPYC) | B/op | allocs/op |
|---|---|---|---|---|
| encoding/json | 4.29 µs | 11.96 µs | 3.03 KiB | 47 |
| encoding/json/v2 | 3.93 µs | 10.91 µs | 3.03 KiB | 47 |
| jsoniter | 4.35 µs | 12.83 µs | 4.93 KiB | 109 |
| sonic | 3.88 µs | 10.03 µs | 6.64 KiB (M5) / 4.58 KiB (VPS) | 51 / 49 |
| easyjson | 2.45 µs | 7.39 µs | 2.94 KiB | 46 |
Encode (marshal the struct back to bytes):
| Library | M5 | VPS (EPYC) | B/op | allocs/op |
|---|---|---|---|---|
| encoding/json | 3.10 µs | 9.60 µs | 4.10 KiB | 21 |
| encoding/json/v2 | 3.44 µs | 10.48 µs | 4.10 KiB | 21 |
| jsoniter | 2.04 µs | 6.23 µs | 4.01 KiB | 20 |
| sonic | 3.11 µs | 7.84 µs | 4.13 KiB | 21 |
| easyjson | 1.85 µs | 5.62 µs | 2.50 KiB | 19 |
Three things worth calling out, because they don't match what the libraries' own marketing says:
easyjson wins both directions on this payload, by a clear margin, on both machines. No JIT warmup, no architecture restrictions - just generated code with no reflection at all.
sonic helps, but it's not "10x." On decode it was about 10% faster than the standard library on the M5 and 16% faster on the VPS - a real but modest edge, not the dramatic win its "JIT-compiled, fastest JSON library" positioning implies - and on both machines it used more memory per op than encoding/json (6.64 KiB on the M5, 4.58 KiB on the VPS, vs. 3.03 KiB for both), which tracks with how it represents the decoded map field internally. Why that gap itself differs by machine - roughly double the standard library's B/op on the M5, about 51% more on the VPS, for the same payload - isn't something this benchmark explains; noted honestly as measured, not accounted for.
jsoniter's decode result depends on which machine you ask. On the VPS it's consistently slower than the standard library across repeated runs (~12.8µs vs ~12.0µs, roughly 7%) - real, if more modest than its marketing implies. On the M5, repeated runs put it statistically indistinguishable from the standard library (4.35µs vs 4.29µs) - a gap that doesn't survive re-measurement, even though an earlier pass of this benchmark had recorded jsoniter as clearly slower there too. Both machines agree on one thing regardless of timing: jsoniter allocates more than twice as much per decode (109 vs 47 allocs), driven by how it walks the map[string]interface{} field - that's a stable, structural difference, not noise. It did help on encode on both machines (roughly 1.5-2.1x). The lesson isn't just "jsoniter is bad" or "jsoniter is fine" - it's that a single-machine, single-run decode number for a reflection-heavy path can simply stop reproducing, and the only way to know if a claim holds for your workload is to benchmark it yourself, more than once, on your actual target.
jsoniter: A Drop-in API, Not a Guaranteed Speedup
Jsoniter is still the easiest first thing to try - it's a drop-in replacement for encoding/json with an identical API, so migration is a one-line import change and nothing else in your code has to move. Whether it's actually faster for your workload is a question the benchmark above answers you shouldn't skip: on encode it won clearly here, on both machines; on decode it won on the VPS and was a wash on the M5.
It maintains close compatibility with standard library behavior, including error messages and edge cases, and offers configuration knobs for cases where you can trade correctness for speed deliberately.
Disabling HTML escaping saves CPU cycles when you know your data is safe. Float precision reduction works well for financial data where 6 digits are sufficient. Custom decoders enable zero-copy operations for specific types, which helps most on string-heavy payloads where copying dominates.
import jsoniter "github.com/json-iterator/go"
var json = jsoniter.ConfigCompatibleWithStandardLibrary
// Drop-in: same call, same error semantics.
// Whether it's faster for YOUR struct is an empirical question - see the
// benchmark above before assuming it is.
err := json.Unmarshal(data, &order)
// Advanced performance configuration
var fastJson = jsoniter.Config{
EscapeHTML: false,
MarshalFloatWith6Digits: true,
ObjectFieldMustBeSimpleString: true,
}.Froze()
// Register for all string fields
jsoniter.RegisterTypeDecoderFunc("string",
func(typ reflect2.Type) jsoniter.ValDecoder {
return StringDecoder{}
})
easyjson: When Every Microsecond Counts
Code generation eliminates reflection entirely:
# Install
go get -u github.com/mailru/easyjson/...
# Generate
//go:generate easyjson -all order.go
// Before generation
type Order struct {
ID string `json:"id"`
UserID int64 `json:"user_id"`
// ... fields
}
// After generation - you get these methods
func (o *Order) MarshalJSON() ([]byte, error)
func (o *Order) UnmarshalJSON(data []byte) error
func (o *Order) MarshalEasyJSON(w *jwriter.Writer)
func (o *Order) UnmarshalEasyJSON(l *jlexer.Lexer)
// Usage - fewer allocations than encoding/json, not zero (46 allocs/op
// measured above for this payload; see the benchmark table)
order := &Order{}
lexer := &jlexer.Lexer{Data: jsonBytes}
order.UnmarshalEasyJSON(lexer)
easyjson Pool Pattern
var orderPool = sync.Pool{
New: func() interface{} {
return &Order{}
},
}
var lexerPool = sync.Pool{
New: func() interface{} {
return &jlexer.Lexer{}
},
}
func ParseOrder(data []byte) (*Order, error) {
order := orderPool.Get().(*Order)
lexer := lexerPool.Get().(*jlexer.Lexer)
lexer.Data = data
order.UnmarshalEasyJSON(lexer)
err := lexer.Error()
lexer.Data = nil // Important: clear reference
lexerPool.Put(lexer)
if err != nil {
orderPool.Put(order)
return nil, err
}
return order, nil
}
// Return to pool after use
defer orderPool.Put(order)
sonic: When CPU Is the Bottleneck
ByteDance's sonic uses JIT compilation, which is a real architectural difference from every other library here - it generates and compiles machine code tailored to the types it's handling at runtime, rather than interpreting a schema through reflection (the actual code-generation and dispatch paths inside sonic are considerably more involved than that one-line summary - worth reading its internals before relying on assumptions about exactly how it specializes per type). On this benchmark it was faster than encoding/json on decode on both machines; on encode it won on the VPS but was a statistical wash against the standard library on the M5 (3.11µs vs 3.10µs - within the run-to-run spread, not a real difference). easyjson still beat it in every case that did show a real gap. This benchmark doesn't establish where sonic's advantage peaks - a different payload shape (deeper nesting, more polymorphic fields, larger arrays) could change the ranking substantially in either direction. Don't extrapolate from one payload shape to yours; benchmark it.
import "github.com/bytedance/sonic"
// Simple usage
err := sonic.Unmarshal(data, &order)
// Pretouch for JIT warmup (critical for performance)
func init() {
var o Order
sonic.Pretouch(reflect.TypeOf(o))
}
One caveat the benchmark numbers above don't cover: they're steady-state throughput, warmed up before timing starts, so JIT compilation and Pretouch cost aren't included. That's the right number for a long-running service where warmup amortizes away - it's the wrong one for serverless, short-lived workers, or a CLI tool that starts fresh each invocation, where JIT compilation happens on every cold start and needs to be measured separately.
Sonic's Tuning Knobs
// Custom config for maximum speed
var sonicAPI = sonic.Config{
NoQuoteTextMarshaler: true,
NoNullSliceOrMap: true,
UseInt64: true, // No number precision loss
CopyString: false, // DANGEROUS: references input buffer
}.Froze()
// With buffer reuse
buf := sonic.NewBuffer()
defer sonic.FreeBuffer(buf)
err := sonicAPI.NewEncoder(buf).Encode(&order)
jsonBytes := buf.Bytes()
The Hidden Gotchas Nobody Mentions
1. easyjson Silently Ignores Fields You Forgot to Regenerate
This isn't a panic - it's quieter and arguably worse. The generated MarshalEasyJSON/UnmarshalEasyJSON methods only know about the fields that existed when you last ran easyjson -all. Add a field and forget to regenerate, and it's silently dropped on encode and silently ignored on decode - no error, no panic, just data that quietly isn't there. CI that runs go generate and fails on a diff catches this; nothing at runtime will.
type Order struct {
ID string `json:"id"`
NewField string `json:"new_field"` // Silently dropped until regenerated
}
2. sonic Doesn't Support All Architectures
// Only amd64 and arm64. Falls back to encoding/json on others
if runtime.GOARCH != "amd64" && runtime.GOARCH != "arm64" {
return json.Unmarshal(data, v) // ~10-16% slower than sonic on this
// benchmark's payload, not 10x - see the decode table above
}
3. jsoniter's ConfigFastest Relaxes UTF-8 Validation
// ConfigFastest skips UTF-8 validation for speed. Don't use it unless
// you explicitly accept that trade-off - validate separately if your
// data source isn't already guaranteed valid UTF-8:
json := jsoniter.ConfigFastest
if !utf8.Valid(data) {
return errors.New("invalid UTF-8")
}
A Runtime-Generic Fallback Pattern
The benchmark says easyjson wins outright here - but easyjson needs generated code per type, so it doesn't fit a shared utility package meant to marshal arbitrary types without a codegen step for each new one. For that generic case (not the hot-path types you'd generate easyjson code for individually), here's a pattern that picks the best runtime-generic option and degrades safely on architectures sonic doesn't support - both branches use their library's standard-compatible config, deliberately, so switching architectures doesn't silently change validation behavior the way mixing sonic.ConfigFastest with jsoniter.ConfigCompatibleWithStandardLibrary would (see the ConfigFastest UTF-8 warning above - the same risk applies to sonic's fastest config, and a fallback path shouldn't have looser guarantees than the primary one):
package jsonutil
import (
"github.com/bytedance/sonic"
jsoniter "github.com/json-iterator/go"
"runtime"
)
var (
jsonAPI API
)
type API interface {
Marshal(v interface{}) ([]byte, error)
Unmarshal(data []byte, v interface{}) error
}
func init() {
// Use sonic on supported platforms - ConfigDefault, not
// ConfigFastest, so both branches validate UTF-8 the same way
if runtime.GOARCH == "amd64" || runtime.GOARCH == "arm64" {
jsonAPI = sonicWrapper{sonic.ConfigDefault}
} else {
// Fallback to jsoniter
jsonAPI = jsoniter.ConfigCompatibleWithStandardLibrary
}
}
// Wrapper to match interface
type sonicWrapper struct {
api sonic.API
}
func (s sonicWrapper) Marshal(v interface{}) ([]byte, error) {
return s.api.Marshal(v)
}
func (s sonicWrapper) Unmarshal(data []byte, v interface{}) error {
return s.api.Unmarshal(data, v)
}
// Public API
func Marshal(v interface{}) ([]byte, error) {
return jsonAPI.Marshal(v)
}
func Unmarshal(data []byte, v interface{}) error {
return jsonAPI.Unmarshal(data, v)
}
What This Means Under Load
The B/op and allocs/op columns in the tables above are per single decode or encode call. Allocation count alone would have led to the wrong conclusion here: encoding/json's 47 allocs/decode and easyjson's 46 are basically the same number, yet easyjson decoded in roughly half the CPU time (4.29µs vs 2.45µs on the M5, 11.96µs vs 7.39µs on the VPS) - the win is in per-allocation cost and reflection overhead, not allocation count. Allocations still matter for GC pressure, and that's where jsoniter's 109 vs everyone else's ~47-51 is the number to flag: more than double the garbage per request, at 50M requests/day, for a library marketed as a speed upgrade. That's the concrete cost of picking a library off its README instead of your own benchmark.
To see this on your own service: go tool pprof -alloc_objects http://your-service/debug/pprof/heap under real production load, not a synthetic benchmark, is what tells you whether JSON is actually where your allocations are going before you spend time switching libraries.
Only Need a Few Fields? Skip the Struct Entirely
Every library above fully decodes the payload into a Go struct, because that's what the benchmark asked them to do. If your actual need is narrower - read 2-3 fields out of a payload you otherwise don't care about - that's a different problem, and gjson solves it by walking the raw bytes for a path match and never building a struct at all:
import "github.com/tidwall/gjson"
id := gjson.GetBytes(data, "id").String()
total := gjson.GetBytes(data, "total").Float()
On the same 2KB payload, reading just id and total this way took 345ns/2 allocs on the M5 and 883ns/2 allocs on the VPS - against 4.41µs/47 allocs and 12.52µs/47 allocs for a full encoding/json decode of the same two fields. That's a genuine 13-14x, because it's doing genuinely less work: no struct allocation, no reflection, no decoding the 12-item slice or the map you're not going to read. It stops being a win the moment you need most of the payload's fields, or need to walk it more than once - at that point you're paying gjson's per-lookup path-parse cost repeatedly instead of decoding once.
Comparing the Options
One clarification before the table: as of Go 1.27, encoding/json - the same json.Marshal/json.Unmarshal everyone already calls - is backed by the v2 implementation internally, by default. It's not an opt-in you switch to; it's already what "StdJSON" in the tables above measured. Per the Go 1.27 release notes: "Marshaling and unmarshaling behavior is preserved, but the exact text of error messages may differ," marshal performance is "broadly at parity" with the old v1 implementation, and unmarshal is "significantly faster." GOEXPERIMENT=nojsonv2 is the escape hatch back to the old v1-only implementation, not a flag to turn v2 on - and that opt-out is expected to be removed in a future release.
More precisely: the legacy encoding/json API delegates to encoding/json/v2 internally through compatibility options that preserve v1's semantics, while encoding/json/v2's native API skips that shim and also defaults to stricter behavior - it rejects invalid UTF-8 in JSON strings and duplicate names within a JSON object, checks v1 doesn't do. So "StdJSON" and "JSONv2" in this benchmark are the same underlying engine, not old vs. new - which is why the allocation counts came out identical (47 decode, 21 encode; there's one implementation doing the allocating). The CPU time didn't come out identical, though: JSONv2 decoded about 8-9% faster on both machines (3.93µs vs 4.29µs on the M5, 10.91µs vs 11.96µs on the VPS, reproducible across repeated runs) - consistent with the release notes' "significantly faster" unmarshal claim, if more modest than "significantly" might suggest for this payload - but encoded about 9-11% slower on both (3.44µs vs 3.10µs on the M5, 10.48µs vs 9.60µs on the VPS), which doesn't obviously match "broadly at parity." That's a real, repeatable measurement on this payload, not explained any further here - if you're choosing between the two APIs specifically for encode-heavy CPU cost, benchmark your own payload rather than assuming parity. The new package is worth reaching for when you want its extra options (RejectUnknownMembers, Deterministic, and others v1 doesn't expose) regardless of which way this particular measurement goes.
| encoding/json | encoding/json/v2 | jsoniter | easyjson | sonic | gjson | |
|---|---|---|---|---|---|---|
| How it works | Reflection, v2 engine (Go 1.27+) | Same engine, v2 API | Reflection (optimized) | Codegen | JIT per type | Byte scan, no struct |
| Full struct decode | Yes | Yes | Yes | Yes | Yes | No - path reads only |
| On this benchmark | Baseline | ~8-9% faster decode, ~9-11% slower encode (same engine, different API path) | Decode: slower on VPS, no diff on M5. Faster encode, both machines | Fastest full decode/encode | ~10-16% faster decode, both machines; faster encode on VPS only | 13-14x for 2 fields only |
| Portability | Any platform | Any platform | Any platform | Any platform | amd64/arm64 only | Any platform |
| Extra step | None | None | None | go generate | None | None |
| Best use case | Default choice on Go 1.27+ | Same default, when you need its extra options | Encode-heavy, standard-library-compatible | Stable schema, hot path | Large/irregular payloads, control your deploy target | Read a few fields, skip the rest |
The Brutal Truth
Most apps don't need this. If you're doing <100 req/s, stick with encoding/json. Profile first.
But if JSON parsing shows up in your profiles as a measurable CPU or allocation cost?
Then these techniques can cut that specific cost substantially - CPU cost, allocation count, and GC pressure don't move together, though, as easyjson's own numbers above show (46 allocs vs the standard library's 47, nearly identical, alongside roughly half the CPU time): picking a library changes whichever of those your profile actually flagged, not automatically all three.
Action Items
- Profile your app first:
go tool pprof -alloc_objectsunder real load. Don't optimize what your profile doesn't flag. - If JSON allocations show up, benchmark your payload against encoding/json, jsoniter, easyjson, and sonic - not this article's numbers. Payload shape changes the ranking, as jsoniter's decode result here shows.
- If you only need a handful of fields out of a large payload, check gjson before reaching for a full decoder at all.
- easyjson's speed comes with a maintenance cost: a regeneration step your CI needs to enforce. Budget for that before adopting it.
- sonic is amd64/arm64 only - keep a fallback path if you deploy anywhere else.
Final tip: we kept encoding/json for low-traffic internal APIs and only switched libraries on the endpoints where profiling actually showed JSON as a bottleneck. Start surgical, expand based on data - not based on which library's README has the biggest number in it.