Streams and data pipelines
A stream describes a pipeline over data. It does not store elements.
Pipeline shape
- Source: collection, array, file lines, generator.
- Intermediate operations:
filter,map,flatMap,sorted,distinct. - Terminal operation:
toList,collect,reduce,count,findFirst.
Intermediate operations are lazy. Nothing runs until a terminal operation needs results.
Stateless behavior
Stream functions should be non-interfering and normally stateless.
Mutating an external list in forEach creates hidden
coupling and breaks parallel safety. Prefer collectors or immutable
transformations.
Mapping and flattening
map turns each element into one result.
flatMap turns each element into zero or more results and
flattens them. mapMulti can emit multiple results without
creating an intermediate stream per input.
Reduction and collectors
A reduction combines elements with an identity and associative
accumulator. For parallel execution, the operation must behave
associatively. Collectors build lists, maps, groups, partitions,
summaries, and custom containers. toMap needs a merge rule
when keys may repeat.
Parallel streams
Parallel does not mean faster. Work uses a common pool by default and depends on data size, splitting, CPU cost, order, blocking, and contention. Avoid parallel streams for blocking I/O and measure CPU pipelines before adopting.
Reuse and side effects
A stream is single-use. Terminal operation closes the pipeline. Streams backed by I/O should be closed with try-with-resources.
Feynman check
A stream is a conveyor belt. Filters remove parcels, maps relabel them, flatMap opens boxes into several parcels, and a collector packs the result.