> built 2026-09-22 14:10 UTC from 85e0943 (master) · recode 0.1.40. Details: build_info.json

# index.html.md

<!-- generated by epythet -->

# recode

Make codecs for fixed size structured chunks serialization and deserialization of
sequences, tabular data, and time-series.

To install:	`pip install recode`

[Docmentation](https://i2mint.github.io/recode/)

Make codecs for fixed size structured chunks serialization and deserialization of
sequences, tabular data, and time-series.

The easiest and bigest bang for your buck is `mk_codec`

```python
>>> from recode import mk_codec
>>> encoder, decoder = mk_codec()
```

`encoder` will encode a list (or any iterable) of numbers into bytes

```python
>>> b = encoder([0, -3, 3.14])
>>> b
b'\x00\x00\x00\x00\x00\x00\x00\x00\x00\x00\x00\x00\x00\x00\x08\xc0\x1f\x85\xebQ\xb8\x1e\t@'

```

`decoder` will decode those bytes to get you back your numbers

```python
>>> decoder(b)
[0.0, -3.0, 3.14]
```

There’s only really one argument you need to know about in `mk_codec`.
The first argument, called `chk_format`, which is a string of characters from
the “Format” column of the python
[format characters](https://docs.python.org/3/library/struct.html#format-characters)

The one we’ve just been through is in fact

```python
>>> encoder, decoder = mk_codec('d')
```

That is, it will expect that your data is a list of numbers, and they’ll be encoded
with the ‘d’ format character, that is 8-bytes doubles.
That default is good because it gives you a lot of room, but if you knew that you
would only be dealing with 2-byte integers (as in most WAV audio waveforms),
you would have chosen `h`:

```python
>>> encoder, decoder = mk_codec('h')
```

What about those channels?
Well, some times you need to encode/decode multi-channel streams, such as:

```python
>>> multi_channel_stream = [[3, -1], [4, -1], [5, -9]]
```

Say, for example, if you were dealing with stereo waveform
(with the standard PCM_16 format), you’d do it this way:

```python
>>> encoder, decoder = mk_codec('hh')
>>> pcm_bytes = encoder(multi_channel_stream)
>>> pcm_bytes
b'\x03\x00\xff\xff\x04\x00\xff\xff\x05\x00\xf7\xff'
>>> decoder(pcm_bytes)
[(3, -1), (4, -1), (5, -9)]
```

The `n_channels` and `chk_size_bytes` arguments are there if you want to assert
that your number of channels and chunk size are what you expect.
Again, these are just for verification, because we know how easy it is to
misspecify the `chk_format`, and how hard it can be to notice that we did.

It is advised to use these in any production code, for the sanity of everyone!

```python
>>> mk_codec('hhh', n_channels=2)
Traceback (most recent call last):
  ...
AssertionError: You said there'd be 2 channels, but I inferred 3
>>> mk_codec('hhh', chk_size_bytes=3)
Traceback (most recent call last):
  ...
AssertionError: The given chk_size_bytes 3 did not match the inferred (from chk_format) 6
```

Finally, so far we’ve done it this way:

```python
>>> encoder, decoder = mk_codec('hHifd')
```

But see that what’s actually returned is a NAMED tuple, which means that you can
can also get one object that will have `.encode` and `.decode` properties:

```python
>>> codec = mk_codec('hHifd')
>>> to_encode = [[1, 2, 3, 4, 5], [6, 7, 8, 9, 10]]
>>> encoded = codec.encode(to_encode)
>>> decoded = codec.decode(encoded)
>>> decoded
[(1, 2, 3, 4.0, 5.0), (6, 7, 8, 9.0, 10.0)]
```

And you can checkout the properties of your encoder and decoder (they
should be the same)

```python
>>> codec.encode.chk_format
'hHifd'
>>> codec.encode.n_channels
5
>>> codec.encode.chk_size_bytes
24
```

# Further functionality and under-the-hood peeps

This section shows various examples of recode and it how it can be used with:

- Single channel numerical streams
- Multi-channel numerical streams
- DataFrames
- Iterators

```python
from recode import (ChunkedEncoder, 
                    ChunkedDecoder, 
                    MetaEncoder, 
                    MetaDecoder, 
                    IterativeDecoder, 
                    StructCodecSpecs, 
                    specs_from_frames,
                    frame_to_meta,
                    meta_to_frame)
```

## Quick run through

First define the frame you want to encode

```python
frame = [1,2,3]
```

Next define the specifications for that encoding

```python
specs = StructCodecSpecs(chk_format='h')
specs
```

```none
StructCodecSpecs(chk_format='h', n_channels=1, chk_size_bytes=2)
```

Next define an encoder and encode your frame to bytes

```python
encoder = ChunkedEncoder(frame_to_chk=specs.frame_to_chk)
b = encoder(frame)
b
```

```none
b'\x01\x00\x02\x00\x03\x00'
```

Once you need your original frame again, define a decoder and decode your frame

```python
decoder = ChunkedDecoder(specs.chk_to_frame)
decoded_frames = decoder(b)
decoded_frames
```

```none
[1, 2, 3]
```

## Step-by-step explanation

### Define your StructCodecSpecs

StructCodecSpecs is used to define the specs for making codecs for fixed size structured chunks serialization and deserialization of sequences, tabular data, and time-series. This definition is based on format strings of the [python struct module](https://docs.python.org/3/library/struct.html#format-strings).

There are two ways to define StructCodecSpecs, the first being to explicitly
define it using the StructCodecSpecs class.
StructCodecSpecs takes three arguments: `chk_format`, `n_channels`,
and `chk_size_bytes`.

Only `chk_format` is required,
as `n_channels` and `chk_size_bytes` can be determined based on `chk_format`.
If `n_channels` or  `chk_size_bytes` are given, they will be used to
assert that the values inferred from `chk_format` match.

If this is not the case, then do not provide an argument for `n_channels`
and instead pass a string with the format character matching the data type for
each channel in the frame to `chk_format`.
For example if the first channel
contains integers and the second contains floats, then `chk_format = 'hd'`.

```python
frame = [1,2,3]
specs = StructCodecSpecs(chk_format='h')
specs
```

```none
StructCodecSpecs(chk_format='h', n_channels=1, chk_size_bytes=2)
```

The second way to define StructCodecSpecs is to use `specs_from_frames` which will implictly define StructCodecSpecs based on the frame that is going to be encoded/decoded. This function will return a tuple containing a reconstituted version of the iterator given as an input, and the defined StructCodecSpecs. The first element of the tuple can be ignored if the frame passed is not an iterator.

```python
_, specs = specs_from_frames(frame)
specs
```

```none
StructCodecSpecs(chk_format='h', n_channels=1, chk_size_bytes=2)
```

If frame is an iterator, then redefine frame as the first argument of the tuple so the first element of frame is not lost for encoding.

```python
frame = iter([[1,2], [3,4]])
frame, specs = specs_from_frames(frame)
specs, list(frame)
```

```none
(StructCodecSpecs(chk_format='hh', n_channels=2, chk_size_bytes=4),
 [[1, 2], [3, 4]])
```

### Define your Encoder

Your Encoder will allow you to, you guessed it, encode your frames! There are two Encoders currently defined in recode: `ChunkedEncoder` and `MetaEncoder`.

`ChunkedEncoder` should be your goto encoder for sequences, while `MetaEncoder` works best for tabular data (currently must be in the format of list of dicts).

```python
frame = [1,2,3]
_, specs = specs_from_frames(frame)
encoder = ChunkedEncoder(frame_to_chk=specs.frame_to_chk)
b = encoder(frame)
b
```

```none
b'\x01\x00\x02\x00\x03\x00'
```

When using a `MetaEncoder`, an extra argument named frame_to_meta is required, which can be easily imported from recode!

```python
frame = [{'foo': 1, 'bar': 1}, {'foo': 2, 'bar': 2}, {'foo': 3, 'bar': 4}]
_, specs = specs_from_frames(frame)
encoder = MetaEncoder(frame_to_chk=specs.frame_to_chk, frame_to_meta=frame_to_meta)
b = encoder(frame)
b
```

```none
b'\x07\x00foo.bar\x01\x00\x01\x00\x02\x00\x02\x00\x03\x00\x04\x00'
```

### Define your Decoder

Next up your Decoder will allow you to decode your encoded bytes. There are three Encoders currently defined in recode: `ChunkedDecoder`, `IterativeDecoder`, and `MetaDecoder`.

Either `ChunkedDecoder` or `IterativeDecoder` will work well for sequences, with the only difference being that `IterativeDecoder` returns an iterator of decoded chunks while `ChunkedDecoder` returns the whole list of decoded chunks. `MetaDecoder` works best for tabular data encoded with `MetaEncoder`.

```python
frame = [1,2,3]
_, specs = specs_from_frames(frame)
encoder = ChunkedEncoder(frame_to_chk=specs.frame_to_chk)
decoder = ChunkedDecoder(specs.chk_to_frame)
b = encoder(frame)
decoder(b)
```

```none
[1, 2, 3]
```

As is shown in the following example, an `IterativeDecoder` will return an unpack_iterator.

```python
frame = [[1,1],[2,2]]
_, specs = specs_from_frames(frame)
encoder = ChunkedEncoder(frame_to_chk=specs.frame_to_chk)
decoder = IterativeDecoder(chk_to_frame=specs.chk_to_frame)
b = encoder(frame)
iter_frames = decoder(b)
print(type(iter_frames))
next(iter_frames), next(iter_frames)
```

```none
<class 'unpack_iterator'>
((1, 1), (2, 2))
```

When using a `MetaDecoder`, an extra argument named meta_to_frame is required, which can be easily imported from recode!

```python
frame = [{'foo': 1, 'bar': 1}, {'foo': 2, 'bar': 2}, {'foo': 3, 'bar': 4}]
_, specs = specs_from_frames(frame)
encoder = MetaEncoder(frame_to_chk=specs.frame_to_chk, frame_to_meta=frame_to_meta)
decoder = MetaDecoder(chk_to_frame=specs.chk_to_frame, meta_to_frame=meta_to_frame)
b = encoder(frame)
decoded_frames = decoder(b)
decoded_frames
```

```none
[{'foo': 1, 'bar': 1}, {'foo': 2, 'bar': 2}, {'foo': 3, 'bar': 4}]
```

## Encoding a waveform

### Create a synthetic waveform using `hum`

```python
from hum.gen.sine_mix import freq_based_stationary_wf
import matplotlib.pyplot as plt
import numpy as np

DFLT_N_SAMPLES = 21 * 2048
DFLT_SR = 44100
```

```python
wf_mix = freq_based_stationary_wf(freqs=(200, 400, 600, 800), weights=None,
                             n_samples = DFLT_N_SAMPLES, sr = DFLT_SR)
plt.plot(wf_mix[:300]);
```

![png](https://raw.githubusercontent.com/i2mint/recode/master/notebooks/recode_demo_files/recode_demo_35_0.png)

```python
wf_mix
```

```none
array([0.        , 0.07114157, 0.14170624, ..., 0.49480205, 0.54853009,
       0.5979443 ])
```

### Encode the wf

```python
specs = StructCodecSpecs('d')
encoder = ChunkedEncoder(frame_to_chk=specs.frame_to_chk)
decoder = ChunkedDecoder(chk_to_frame=specs.chk_to_frame)
b = encoder(wf_mix)
b[:100]
```

```none
b'\x00\x00\x00\x00\x00\x00\x00\x00\x1b:\x11\x8dU6\xb2?\xb8!\xa2\x15n#\xc2?\xe6I\xedz\x15\x06\xcb?L;\xe6\xfah\xd8\xd1?}\x8c\x9b\x9a\xf5\x08\xd6?D\xcc\xf1\x897\x0c\xda?\xe3\xc5\xbeb/\xda\xdd?DE\xb6K\xb6\xb5\xe0?\x82|\xfcz\x90\\\xe2?\xae3:$\x99\xde\xe3?\x1e\x95\xbe\xef$9\xe5?\x81\xd7\xdc\xec'
```

### Decode and compare to wf_mix

```python
decoded_wf_mix = decoder(b)
plt.plot(decoded_wf_mix[:300]);
```

![png](https://raw.githubusercontent.com/i2mint/recode/master/notebooks/recode_demo_files/recode_demo_40_0.png)

```python
np.all(decoded_wf_mix == wf_mix)
```

```none
True
```

## Encoding a pandas dataframe

```python
import pandas as pd
```

### Create/import your dataframe

```python
df = pd.DataFrame(data = [[1,2,3],[4,5,6]], columns = ['foo', 'bar', 'set'])
df
```

<div>
<table border="1" class="dataframe">
  <thead>
    <tr style="text-align: right;">
      <th></th>
      <th>foo</th>
      <th>bar</th>
      <th>set</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <th>0</th>
      <td>1</td>
      <td>2</td>
      <td>3</td>
    </tr>
    <tr>
      <th>1</th>
      <td>4</td>
      <td>5</td>
      <td>6</td>
    </tr>
  </tbody>
</table>
</div>

### Prep the dataframe for encoding

```python
frame = df.to_dict('records')
frame
```

```none
[{'foo': 1, 'bar': 2, 'set': 3}, {'foo': 4, 'bar': 5, 'set': 6}]
```

### Encode the list of dicts

```python
_, specs = specs_from_frames(frame)
encoder = MetaEncoder(frame_to_chk=specs.frame_to_chk, frame_to_meta=frame_to_meta)
decoder = MetaDecoder(chk_to_frame=specs.chk_to_frame, meta_to_frame=meta_to_frame)
b = encoder(frame)
b
```

```none
b'\x0b\x00foo.bar.set\x01\x00\x02\x00\x03\x00\x04\x00\x05\x00\x06\x00'
```

```python
decoded_frame = decoder(b)
decoded_frame
```

```none
[{'foo': 1, 'bar': 2, 'set': 3}, {'foo': 4, 'bar': 5, 'set': 6}]
```

```python
decoded_df = pd.DataFrame(decoded_frame)
print(np.all(decoded_df == df))
decoded_df
```

```none
True
```

<div>
<table border="1" class="dataframe">
  <thead>
    <tr style="text-align: right;">
      <th></th>
      <th>foo</th>
      <th>bar</th>
      <th>set</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <th>0</th>
      <td>1</td>
      <td>2</td>
      <td>3</td>
    </tr>
    <tr>
      <th>1</th>
      <td>4</td>
      <td>5</td>
      <td>6</td>
    </tr>
  </tbody>
</table>
</div>

## Miscellaneous information

### Byte order, Size, and Alignment

This table provides the characters associated with Byte order, Size, and Alignment. If one of these characters is not given as the first character of the format string, then `@` will be assumed. More information about Byte order, Size, and Alignment can be found [here](https://docs.python.org/3/library/struct.html#byte-order-size-and-alignment).

| Character   | Byte order            | Size     | Alignment   |
|-------------|-----------------------|----------|-------------|
| @           | native                | native   | native      |
| =           | native                | standard | none        |
| <           | little-endian         | standard | none        |
| >           | big-endian            | standard | none        |
| !           | network (=big-endian) | standard | none        |

# Full Examples

## Single channel numerical stream

```python
from recode import StructCodecSpecs, ChunkedEncoder, ChunkedDecoder
specs = StructCodecSpecs(chk_format='h')
encoder = ChunkedEncoder(frame_to_chk=specs.frame_to_chk)
decoder = ChunkedDecoder(chk_to_frame=specs.chk_to_frame)
frames = [1, 2, 3]
b = encoder(frames)
assert b == b'\x01\x00\x02\x00\x03\x00'
decoded_frames = decoder(b)
assert decoded_frames == frames
```

## Multi-channel numerical stream

```python
from recode import StructCodecSpecs, ChunkedEncoder, ChunkedDecoder
specs = StructCodecSpecs(chk_format='@hh', n_channels = 2)
encoder = ChunkedEncoder(frame_to_chk=specs.frame_to_chk)
decoder = ChunkedDecoder(chk_to_frame=specs.chk_to_frame)
frames = [(1, 2), (3, 4), (5, 6)]
b = encoder(frames)
assert b == b'\x01\x00\x02\x00\x03\x00\x04\x00\x05\x00\x06\x00'
decoded_frames = decoder(b)
assert decoded_frames == frames
```

## Iterative decoder

```python
from recode import StructCodecSpecs, ChunkedEncoder, IterativeDecoder
specs = StructCodecSpecs(chk_format = 'hdhd')
encoder = ChunkedEncoder(frame_to_chk = specs.frame_to_chk)
decoder = IterativeDecoder(chk_to_frame = specs.chk_to_frame)
frames = [(1,1.1,1,1.1),(2,2.2,2,2.2),(3,3.3,3,3.3)]
b = encoder(frames)
iter_frames = decoder(b)
assert next(iter_frames) == frames[0]
next(iter_frames)
```

```none
(2, 2.2, 2, 2.2)
```

## DataFrame (as list of dicts) using MetaEncoder/MetaDecoder

```python
from recode import StructCodecSpecs, MetaEncoder, MetaDecoder, frame_to_meta, meta_to_frame
data = [{'foo': 1.1, 'bar': 2.2}, {'foo': 513.23, 'bar': 456.1}, {'foo': 32.0, 'bar': 6.7}]
specs = StructCodecSpecs(chk_format='dd', n_channels = 2)
encoder = MetaEncoder(frame_to_chk = specs.frame_to_chk, frame_to_meta = frame_to_meta)
decoder = MetaDecoder(chk_to_frame = specs.chk_to_frame, meta_to_frame = meta_to_frame)
b = encoder(data)
assert decoder(b) == data
```

## Implicitly define codec specs based on frames

```python
from recode import specs_from_frames, ChunkedEncoder, ChunkedDecoder
frames = [1,2,3]
_, specs = specs_from_frames(frames)
encoder = ChunkedEncoder(frame_to_chk = specs.frame_to_chk)
decoder = ChunkedDecoder(chk_to_frame=specs.chk_to_frame)
b = encoder(frames)
assert b == b'\x01\x00\x02\x00\x03\x00'
decoded_frames = decoder(b)
decoded_frames
```

```none
[1, 2, 3]
```

## Implicit definition of iterator

```python
from recode import specs_from_frames, ChunkedEncoder, IterativeDecoder
frames = iter([[1.1,2.2],[3.3,4.4]])
frames, specs = specs_from_frames(frames)
encoder = ChunkedEncoder(frame_to_chk = specs.frame_to_chk)
decoder = ChunkedDecoder(chk_to_frame=specs.chk_to_frame)
b = encoder(frames)
decoded_frames = list(decoder(b))
assert decoded_frames == [(1.1,2.2),(3.3,4.4)]
```

# More

## Example of recode functionality to read and write audio files

In the below example we can see that the functionality of reading and writing audio to bytes is replicated in `recode`.
An example of a waveform audio file is first created using `mk_wf`.
Then, that example wave form is converted to bytes using `BytesIO` and `soundfile.write`.
Then, that same example wave form is converted to bytes using `recode` functionality.
Finally it can be seen through the assertions that `soundfile`’s read/write and `recode`’s encode/decode provide
the same functionality for audio files (aside from the header in `soundfile` as a result of the .wav format).

```python
import soundfile as sf
from io import BytesIO
from enum import Enum
import numpy as np
from recode import ChunkedDecoder, ChunkToFrame, ChunkedEncoder, StructCodecSpecs

class Kind(Enum):
    random = 'random'
    increasing = 'increasing'

def mk_wf(n_samples=2048, kind: Kind=Kind.random, **kwargs):
    if kind == Kind.random:
        dtype_str = kwargs.get('num_type', 'int16')
        if dtype_str.startswith('int'):
            low = kwargs.get('low', int(-2**15))
            high = kwargs.get('high', int(2**15-1))
            wf = np.random.randint(low=low, high=high, size=n_samples, dtype=dtype_str)
        else:
            raise TypeError("Don't know how to handle this case")
    else:
        raise TypeError("Don't know how to handle this case")
    return wf

wf = mk_wf(25)
sr = 10000

file = BytesIO()
sf.write(file, wf, sr, format = 'WAV')
file.seek(0)
b = file.read()

specs = StructCodecSpecs('h')
encoder = ChunkedEncoder(frame_to_chk=specs.frame_to_chk)
decoder = ChunkedDecoder(chk_size_bytes=specs.chk_size_bytes, 
                         chk_to_frame=specs.chk_to_frame, 
                         n_channels=specs.n_channels)
d = encoder(wf)

assert d == b[44:]

sf_read = sf.read(BytesIO(b), dtype='int16')
recode_read = decoder(d)

assert np.all(sf_read[0] == recode_read)
assert np.all(recode_read == wf)
assert np.all(sf_read[0] == wf)
```

<p class="epythet-aggregates">This documentation as a single file: <a href="recode.md">recode.md</a> (Markdown, for agents).</p>


# _autosummary/recode.audio.html.md

# recode.audio

Encoding audio

This module illustrates how one can use recode to make audio codecs.

```pycon
>>> wav_bytes = encode_wav_bytes([1, 2, 3], 42)
>>> wav_bytes
b'RIFF*\x00\x00\x00WAVEfmt \x10\x00\x00\x00\x01\x00\x01\x00*\x00\x00\x00T\x00\x00\x00\x02\x00\x10\x00data\x06\x00\x00\x00\x01\x00\x02\x00\x03\x00'
>>> wf, sr = decode_wav_bytes(wav_bytes)
>>> sr
42
>>> wf
[1, 2, 3]
```

The wav codecs are based on pcm codecs along with wav header codecs
(i.e. parsing and generation – using the builtin `wave` package).

Make pcm encoders and decoders:

```pycon
>>> encode, decode = mk_pcm_audio_codec('int16')
>>> encoded = encode([1, 2, 3])
>>> encoded
b'\x01\x00\x02\x00\x03\x00'
>>> decode(encoded)
[1, 2, 3]
```

Or just encode directly:

```pycon
>>> encode_pcm_bytes([1, 2, 3])
b'\x01\x00\x02\x00\x03\x00'
```

Or decode directly:

```pycon
>>> encode_pcm_bytes([1, 2, 3])
b'\x01\x00\x02\x00\x03\x00'
```

### Functions

| [`decode_pcm_bytes`](_autosummary/recode.audio.html.md#recode.audio.decode_pcm_bytes)(pcm_bytes[, width, n_channels])   |                                                                                   |
|-----------------------------------------------------------------------------------------------------|-----------------------------------------------------------------------------------|
| [`decode_wav_bytes`](_autosummary/recode.audio.html.md#recode.audio.decode_wav_bytes)(wav_bytes, \*[, ...])             | Decode WAV bytes into a `(waveform, sample_rate)` pair.                           |
| [`decode_wav_header_bytes`](_autosummary/recode.audio.html.md#recode.audio.decode_wav_header_bytes)(wav_header_bytes)          | Get a dict of params decoded from a wav header                                    |
| [`encode_pcm_bytes`](_autosummary/recode.audio.html.md#recode.audio.encode_pcm_bytes)(wf[, width, n_channels])          | Encode waveform (e.g. list of numbers) into PCM bytes.                            |
| [`encode_wav_bytes`](_autosummary/recode.audio.html.md#recode.audio.encode_wav_bytes)(wf, sr[, width_bytes, ...])       | Encode waveform (e.g. list of numbers) into PCM bytes with WAV header.            |
| [`encode_wav_header_bytes`](_autosummary/recode.audio.html.md#recode.audio.encode_wav_header_bytes)(sr, width_bytes, \*)       | Make a WAV header from given parameters.                                          |
| [`extract_wav_header_from_file`](_autosummary/recode.audio.html.md#recode.audio.extract_wav_header_from_file)(filepath, \*[, ...])  | Extract the header of a WAV file -- everything before the audio -- from its path. |
| [`header_size_of_wav_bytes`](_autosummary/recode.audio.html.md#recode.audio.header_size_of_wav_bytes)(wav_bytes[, meta])        | Size, in bytes, of everything preceding the audio payload.                        |
| [`mk_pcm_audio_codec`](_autosummary/recode.audio.html.md#recode.audio.mk_pcm_audio_codec)([width, n_channels])            | Make a (encoder, decoder) pair for PCM data with given width and n_channels.      |
| [`num_find_num_type_for`](_autosummary/recode.audio.html.md#recode.audio.num_find_num_type_for)(num[, target_num_sys, ...])  | Find the target_num_sys equivalent of input num checking multiple unit options    |
| [`num_type_for`](_autosummary/recode.audio.html.md#recode.audio.num_type_for)(num[, num_sys, target_num_sys])       | Translate from one (sample width) number type to another.                         |

### Exceptions

| [`ShortWavData`](_autosummary/recode.audio.html.md#recode.audio.ShortWavData)   | The `data` chunk carries fewer bytes than its own header declares.   |
|-----------------------------------------------------------------|----------------------------------------------------------------------|

### *exception* recode.audio.ShortWavData

Bases: [`UserWarning`](https://docs.python.org/3/builtins/exceptions.html#UserWarning)

The `data` chunk carries fewer bytes than its own header declares.

Raised as a warning rather than an error because the audio that *is* present is
still worth decoding – a partially downloaded file, or one written to a stream
whose length was never patched back into the header. What must not happen is for
the shortfall to pass unmentioned, since the caller cannot otherwise tell a
truncated file from a complete one.

### recode.audio.decode_pcm_bytes(pcm_bytes, width=2, n_channels=1)

* **Parameters:**
  * **width** (`Union`[[`str`](https://docs.python.org/3/builtins/stdtypes.html#str), [`int`](https://docs.python.org/3/builtins/functions.html#int)]) – The width of a sample (in bits, bytes, numpy dtype, pyaudio …)
    (Will try to figure it out)
  * **n_channels** ([`int`](https://docs.python.org/3/builtins/functions.html#int)) – Number of channels
* **Returns:**
  The decoded waveform

```pycon
>>> decode_pcm_bytes(b'\x01\x00\x02\x00\x03\x00')
[1, 2, 3]
```

### recode.audio.decode_wav_bytes(wav_bytes, , eight_bit_unsigned=True)

Decode WAV bytes into a `(waveform, sample_rate)` pair.

* **Parameters:**
  * **wav_bytes** ([`bytes`](https://docs.python.org/3/builtins/stdtypes.html#bytes)) – The bytes of a RIFF/WAVE container holding uncompressed PCM
  * **eight_bit_unsigned** ([`bool`](https://docs.python.org/3/builtins/functions.html#bool)) – Read 8-bit audio as the unsigned PCM the WAV spec
    mandates (0..255 on disk, biased to -128..127 here). Pass `False` for the
    pre-recode#12 behaviour, which read those bytes as signed – silence came back
    as -128 – and so round-tripped with `encode_wav_bytes` but with nothing else.
    Widths of 9 bits and up are signed in the spec and are unaffected either way.
* **Returns:**
  `(wf, sr)` – the decoded waveform and its sample rate
* **Raises:**
  * [**ValueError**](https://docs.python.org/3/builtins/exceptions.html#ValueError) – if `wav_bytes` is not a RIFF/WAVE container with a `data`
    chunk. (Before recode#4 the same inputs raised `AssertionError`, `wave.Error`
    or `EOFError` depending on how they were malformed; they are unified here.)
  * [**ShortWavData**](_autosummary/recode.audio.html.md#recode.audio.ShortWavData) – *warning*, not an exception – see below.

```pycon
>>> wav_bytes = (
...     b'RIFF.\x00\x00\x00WAVEfmt \x10\x00\x00\x00\x01\x00\x01\x00'  # header
...     b'*\x00\x00\x00T\x00\x00\x00\x02\x00\x10\x00data\n\x00\x00\x00'  # header
...     b'\x00\x00\x01\x00\xff\xff\x02\x00\xfe\xff'  # data
... )
>>> wf, sr = decode_wav_bytes(wav_bytes)
>>> wf
[0, 1, -1, 2, -2]
>>> sr
42
```

The `data` chunk is located by walking the RIFF structure, so chunks that sit
*after* the audio – `LIST`/`INFO` metadata, which ffmpeg, Audacity and iTunes
all append – do not shift the waveform:

```pycon
>>> import struct
>>> info = b'INFOISFT' + struct.pack('<I', 6) + b'Lavf58'
>>> with_trailing_metadata = wav_bytes + b'LIST' + struct.pack('<I', len(info)) + info
>>> decode_wav_bytes(with_trailing_metadata)[0]
[0, 1, -1, 2, -2]
```

A file carrying less audio than its header declares decodes to the whole frames
that are actually there, and says so:

```pycon
>>> truncated = wav_bytes[:-4]
>>> import warnings
>>> with warnings.catch_warnings(record=True) as caught:
...     _ = warnings.simplefilter('always')
...     wf, sr = decode_wav_bytes(truncated)
>>> wf
[0, 1, -1]
>>> caught[0].category.__name__
'ShortWavData'
```

8-bit WAV audio is stored *unsigned*, so a byte of 128 is silence rather than full
negative (recode#12). Pass `eight_bit_unsigned=False` to get the old signed reading:

```pycon
>>> eight_bit = encode_wav_bytes([-128, 0, 127], sr=42, width_bytes=1)
>>> eight_bit[44:]
b'\x00\x80\xff'
>>> decode_wav_bytes(eight_bit)[0]
[-128, 0, 127]
>>> decode_wav_bytes(eight_bit, eight_bit_unsigned=False)[0]
[0, -128, -1]
```

### recode.audio.decode_wav_header_bytes(wav_header_bytes)

Get a dict of params decoded from a wav header

For examples, see the `encode_wav_header_bytes` function, it’s inverse.

* **Return type:**
  [`dict`](https://docs.python.org/3/builtins/stdtypes.html#dict)

```pycon
>>> from recode.audio import encode_wav_header_bytes
>>> header_bytes = encode_wav_header_bytes(44100, 2, n_channels=3)
>>> decode_wav_header_bytes(header_bytes)
{'sr': 44100,
 'width_bytes': 2,
 'n_channels': 3,
 'nframes': 0,
 'comptype': None}
```

Stdlib `wave` is the primary reader. A `WAVE_FORMAT_EXTENSIBLE` header, which it
refuses before Python 3.12, is parsed directly instead so the same file decodes on
every supported version – see `_extensible_wav_header_params()`.

### recode.audio.encode_pcm_bytes(wf, width=16, n_channels=1)

Encode waveform (e.g. list of numbers) into PCM bytes.

* **Parameters:**
  * **wf** ([`Iterable`](https://docs.python.org/3/library/collections.abc.html#collections.abc.Iterable)[[`Number`](https://docs.python.org/3/library/numbers.html#numbers.Number)]) – Waveform to encode
  * **width** (`Union`[[`str`](https://docs.python.org/3/builtins/stdtypes.html#str), [`int`](https://docs.python.org/3/builtins/functions.html#int)]) – The width of a sample (in bits, bytes, numpy dtype, pyaudio …)
    (will try to figure it out by itself)
  * **n_channels** ([`int`](https://docs.python.org/3/builtins/functions.html#int)) – Number of channels
* **Returns:**
  The pcm-bytes-encoded waveform

```pycon
>>> encode_pcm_bytes([1, 2, 3])
b'\x01\x00\x02\x00\x03\x00'
```

### recode.audio.encode_wav_bytes(wf, sr, width_bytes=2, n_channels=1, , eight_bit_unsigned=True)

Encode waveform (e.g. list of numbers) into PCM bytes with WAV header.

* **Parameters:**
  * **wf** ([`Iterable`](https://docs.python.org/3/library/collections.abc.html#collections.abc.Iterable)[[`Number`](https://docs.python.org/3/library/numbers.html#numbers.Number)]) – Waveform to encode (iterable of numbers)
  * **sr** ([`int`](https://docs.python.org/3/builtins/functions.html#int)) – Sample rate in Hz
  * **width_bytes** ([`int`](https://docs.python.org/3/builtins/functions.html#int)) – The width of a sample in bytes
  * **n_channels** ([`int`](https://docs.python.org/3/builtins/functions.html#int)) – Number of channels
  * **eight_bit_unsigned** ([`bool`](https://docs.python.org/3/builtins/functions.html#bool)) – Write 8-bit audio as the unsigned PCM the WAV spec
    mandates (samples biased by 128 into 0..255 on disk). Pass `False` for the
    pre-recode#12 behaviour, which wrote them signed – files that recode read
    back correctly and every other tool did not. Widths of 9 bits and up are
    signed in the spec and are unaffected either way.
* **Returns:**
  The complete WAV file bytes (header + data)
* **Return type:**
  [*bytes*](https://docs.python.org/3/builtins/stdtypes.html#bytes)

### Examples

```pycon
>>> wav_bytes = encode_wav_bytes([0, 1, -1, 2, -2], sr=42)
>>> header_bytes, data_bytes = wav_bytes[:44], wav_bytes[44:]
>>> data_bytes
b'\x00\x00\x01\x00\xff\xff\x02\x00\xfe\xff'
```

See that the header bytes can be decoded to get the right information about our waveform:

```pycon
>>> decode_wav_header_bytes(header_bytes)
{'sr': 42, 'width_bytes': 2, 'n_channels': 1, 'nframes': 5, 'comptype': None}
```

See that our wave_bytes can be decoded to get the original waveform and sample rate:

```pycon
>>> decoded_wf, decoded_sr = decode_wav_bytes(wav_bytes)
>>> decoded_wf
[0, 1, -1, 2, -2]
>>> decoded_sr
42
```

### recode.audio.encode_wav_header_bytes(sr, width_bytes, , n_channels=1, nframes=0, comptype=None)

Make a WAV header from given parameters.

* **Parameters:**
  * **sr** ([`int`](https://docs.python.org/3/builtins/functions.html#int)) – The sample rate (i.e. “frame rate” i.e. “chk_rate”)
  * **width_bytes** ([`int`](https://docs.python.org/3/builtins/functions.html#int)) – The “sample width” in bytes
  * **n_channels** ([`int`](https://docs.python.org/3/builtins/functions.html#int)) – Number of channels (default is 1)
  * **nframes** ([`int`](https://docs.python.org/3/builtins/functions.html#int)) – 

    Optional number of frames (default is 0).

    NOTE:
    : If a wav file is to be read correctly, the num of frames (i.e.
      samples/chks) should be exactly the number you’ll actually be writing in the
      wave file.
  * **comptype** – No supported by python’s wave module (yet).
* **Return type:**
  [`bytes`](https://docs.python.org/3/builtins/stdtypes.html#bytes)

```pycon
>>> header_bytes = encode_wav_header_bytes(44100, 2, n_channels=3)
>>> len(header_bytes)
44
>>> header_bytes[:31]
b'RIFF$\x00\x00\x00WAVEfmt \x10\x00\x00\x00\x01\x00\x03\x00D\xac\x00\x00\x98\t\x04'
```

You can decode those params (including those you didn’t specify, but were
defaulted) with the `decode_wav_header_bytes` inverse function.

```pycon
>>> from recode.audio import decode_wav_header_bytes
>>> params = decode_wav_header_bytes(header_bytes)
>>> params
{'sr': 44100,
 'width_bytes': 2,
 'n_channels': 3,
 'nframes': 0,
 'comptype': None}
>>> assert encode_wav_header_bytes(**params) == header_bytes
```

### recode.audio.extract_wav_header_from_file(filepath, , read_size=65536)

Extract the header of a WAV file – everything before the audio – from its path.

Useful for reading the header of a WAV file without having to read the entire file
into memory, which is what you want when WAV files are large and/or numerous: only
as much of the file as the header occupies is read.

The answer is the same one [`header_size_of_wav_bytes()`](_autosummary/recode.audio.html.md#recode.audio.header_size_of_wav_bytes) gives for the same
bytes, because both locate the `data` chunk by walking the RIFF structure rather
than inferring where it must be. This used to compute
`chunk_size + 8 - subchunk2_size`, reading bytes 40-44 as the size of the audio –
true only of a bare 44-byte header. With a `LIST`/`INFO` chunk after the audio (what
ffmpeg, Audacity and iTunes write) it returned audio bytes as header; with one
before, a truncated prefix that parses as nothing at all (recode#4).

* **Parameters:**
  * **filepath** – The path to the WAV file
  * **read_size** ([`int`](https://docs.python.org/3/builtins/functions.html#int)) – How many bytes to read at a time while looking for the audio
* **Returns:**
  The bytes of the WAV file header, i.e. everything preceding the `data`
  chunk’s contents
* **Raises:**
  [**ValueError**](https://docs.python.org/3/builtins/exceptions.html#ValueError) – if the file is not a RIFF/WAVE container with a reachable `data`
  chunk. (It used to answer such files with arbitrary bytes and no complaint.)

### recode.audio.header_size_of_wav_bytes(wav_bytes, meta=None)

Size, in bytes, of everything preceding the audio payload.

That is the offset of the `data` chunk’s contents, found by walking the RIFF
structure (see `_wav_data_chunk()`). For a well-formed file with nothing after
the audio this is the same number the old size-subtraction produced; unlike it, it
stays correct when the file carries trailing metadata or an over-declared `data`
size.

`meta` – an already-decoded header, once passed to save re-parsing it – is
accepted for backwards compatibility and no longer used.

* **Return type:**
  [`int`](https://docs.python.org/3/builtins/functions.html#int)

```pycon
>>> header_size_of_wav_bytes(
...     b'RIFF.\x00\x00\x00WAVEfmt \x10\x00\x00\x00\x01\x00\x01\x00'
...     b'*\x00\x00\x00T\x00\x00\x00\x02\x00\x10\x00data\n\x00\x00\x00'
...     b'\x00\x00\x01\x00\xff\xff\x02\x00\xfe\xff'
... )
44
```

### recode.audio.mk_pcm_audio_codec(width=16, n_channels=1)

Make a (encoder, decoder) pair for PCM data with given width and n_channels.

PCM data is what’s used in the uncompressed raw WAVE formats (such as used in CDs).
See [https://en.wikipedia.org/wiki/Pulse-code_modulation](https://en.wikipedia.org/wiki/Pulse-code_modulation).

* **Parameters:**
  * **width** (`Union`[[`str`](https://docs.python.org/3/builtins/stdtypes.html#str), [`int`](https://docs.python.org/3/builtins/functions.html#int)]) – The width of a sample (in bits, bytes, numpy dtype, pyaudio …)
    (Will try to figure it out)
  * **n_channels** ([`int`](https://docs.python.org/3/builtins/functions.html#int)) – Number of channels
* **Returns:**
  A (encoder, decoder) pair of functions that are inverse of each other

```pycon
>>> encode, decode = mk_pcm_audio_codec('int16')
>>> encoded = encode([1, 2, 3])
>>> encoded
b'\x01\x00\x02\x00\x03\x00'
>>> decode(encoded)
[1, 2, 3]
```

Let’s check over more combinations of width and n_channels that we can decode
what we encode to get back the same thing:

```pycon
>>> wf = [-3, -2, -1, 0, 1, 2, 3]
>>> for width in [16, 2, 'int16', 'paInt16', 'PCM_16', 32, 4, 'int32']:
...     for channel in wf:
...         encode, decode = mk_pcm_audio_codec('int16')
...         encoded = encode(wf)
...         assert isinstance(encoded, bytes)
...         assert decode(encoded) == wf
```

### recode.audio.num_find_num_type_for(num, target_num_sys='struct', num_sys_search_order=('n_bits', 'n_bytes', 'dtype', 'pyaudio', 'soundfile'))

Find the target_num_sys equivalent of input num checking multiple unit options

### recode.audio.num_type_for(num, num_sys='n_bits', target_num_sys='struct')

Translate from one (sample width) number type to another.

* **Parameters:**
  * **num**
  * **num_sys**
  * **target_num_sys**
* **Returns:**

```pycon
>>> num_type_for(16, "n_bits", "soundfile")
'PCM_16'
>>> num_type_for(3, "n_bytes", "soundfile")
'PCM_24'
```

#### TIP
Use with `functools.partial` when you have some fix translation endpoints.

```pycon
>>> from functools import partial
>>> get_dtype_from_n_bytes = partial(
...     num_type_for, num_sys="n_bytes", target_num_sys="dtype"
... )
>>> get_dtype_from_n_bytes(8)
'float64'
```


# _autosummary/recode.base.html.md

# recode.base

Base recode objects

### Functions

| `add_coding_attributes`(to_obj, from_obj)                                                  |                                                                                                                                                                                                                                  |
|--------------------------------------------------------------------------------------------|----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|
| [`frame_to_meta`](_autosummary/recode.base.html.md#recode.base.frame_to_meta)(frame)                      | Defines header for serialization of tabluar data                                                                                                                                                                                 |
| [`meta_to_frame`](_autosummary/recode.base.html.md#recode.base.meta_to_frame)(meta)                       | Deserializes header for deserialization of tabular data                                                                                                                                                                          |
| [`mk_codec`](_autosummary/recode.base.html.md#recode.base.mk_codec)([chk_format, n_channels, ...])   | Enable the definition of codec specs based on format characters of the python struct module ([https://docs.python.org/3/library/struct.html#format-characters](https://docs.python.org/3/library/struct.html#format-characters)) |
| [`mk_encoder_and_decoder`](_autosummary/recode.base.html.md#recode.base.mk_encoder_and_decoder)([chk_format, ...]) | Enable the definition of codec specs based on format characters of the python struct module ([https://docs.python.org/3/library/struct.html#format-characters](https://docs.python.org/3/library/struct.html#format-characters)) |
| [`specs_from_frames`](_autosummary/recode.base.html.md#recode.base.specs_from_frames)(frames)                 | Implicitly defines the codec specs based on the frames to encode/decode.                                                                                                                                                         |

### Classes

| [`ChunkedDecoder`](_autosummary/recode.base.html.md#recode.base.ChunkedDecoder)(chk_to_frame[, chk_format, ...])   | Deserializes numerical streams and sequences serialized by ChunkedEncoder                                                                                                                                                        |
|----------------------------------------------------------------------------------------------------|----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------|
| [`ChunkedEncoder`](_autosummary/recode.base.html.md#recode.base.ChunkedEncoder)(frame_to_chk[, chk_format, ...])   | Serializes numerical streams and sequences                                                                                                                                                                                       |
| [`IterativeDecoder`](_autosummary/recode.base.html.md#recode.base.IterativeDecoder)(chk_to_frame)                    | Creates an iterator of deserialized chunks of numerical streams and sequences serialized by ChunkedEncoder                                                                                                                       |
| [`MetaDecoder`](_autosummary/recode.base.html.md#recode.base.MetaDecoder)(chk_to_frame, meta_to_frame)          | Deserializes tabular data serialized by MetaEncoder                                                                                                                                                                              |
| [`MetaEncoder`](_autosummary/recode.base.html.md#recode.base.MetaEncoder)(frame_to_chk, frame_to_meta)          | Serializes tabular data (must be formatted as list of dicts)                                                                                                                                                                     |
| [`StructCodecSpecs`](_autosummary/recode.base.html.md#recode.base.StructCodecSpecs)([chk_format, n_channels, ...])   | Enable the definition of codec specs based on format characters of the python struct module ([https://docs.python.org/3/library/struct.html#format-characters](https://docs.python.org/3/library/struct.html#format-characters)) |
| [`codec_tuple`](_autosummary/recode.base.html.md#recode.base.codec_tuple)(encode, decode)                       |                                                                                                                                                                                                                                  |

### *class* recode.base.ChunkedDecoder(chk_to_frame, chk_format=None, n_channels=None, chk_size_bytes=None)

Bases: [`Callable`](https://docs.python.org/3/library/collections.abc.html#collections.abc.Callable)[[[`bytes`](https://docs.python.org/3/builtins/stdtypes.html#bytes)], [`Iterable`](https://docs.python.org/3/library/collections.abc.html#collections.abc.Iterable)[[`Any`](https://docs.python.org/3/library/typing.html#typing.Any) | [`Sequence`](https://docs.python.org/3/library/collections.abc.html#collections.abc.Sequence)[[`Any`](https://docs.python.org/3/library/typing.html#typing.Any)]]]

Deserializes numerical streams and sequences serialized by ChunkedEncoder

### *class* recode.base.ChunkedEncoder(frame_to_chk, chk_format=None, n_channels=None, chk_size_bytes=None)

Bases: [`Callable`](https://docs.python.org/3/library/collections.abc.html#collections.abc.Callable)[[[`Iterable`](https://docs.python.org/3/library/collections.abc.html#collections.abc.Iterable)[[`Any`](https://docs.python.org/3/library/typing.html#typing.Any) | [`Sequence`](https://docs.python.org/3/library/collections.abc.html#collections.abc.Sequence)[[`Any`](https://docs.python.org/3/library/typing.html#typing.Any)]]], [`bytes`](https://docs.python.org/3/builtins/stdtypes.html#bytes)]

Serializes numerical streams and sequences

### *class* recode.base.IterativeDecoder(chk_to_frame)

Bases: [`Callable`](https://docs.python.org/3/library/collections.abc.html#collections.abc.Callable)[[[`bytes`](https://docs.python.org/3/builtins/stdtypes.html#bytes)], [`Iterable`](https://docs.python.org/3/library/collections.abc.html#collections.abc.Iterable)[[`Any`](https://docs.python.org/3/library/typing.html#typing.Any) | [`Sequence`](https://docs.python.org/3/library/collections.abc.html#collections.abc.Sequence)[[`Any`](https://docs.python.org/3/library/typing.html#typing.Any)]]]

Creates an iterator of deserialized chunks of numerical streams and sequences serialized
by ChunkedEncoder

### *class* recode.base.MetaDecoder(chk_to_frame, meta_to_frame)

Bases: [`Callable`](https://docs.python.org/3/library/collections.abc.html#collections.abc.Callable)[[[`bytes`](https://docs.python.org/3/builtins/stdtypes.html#bytes)], [`Iterable`](https://docs.python.org/3/library/collections.abc.html#collections.abc.Iterable)[[`Any`](https://docs.python.org/3/library/typing.html#typing.Any) | [`Sequence`](https://docs.python.org/3/library/collections.abc.html#collections.abc.Sequence)[[`Any`](https://docs.python.org/3/library/typing.html#typing.Any)]]]

Deserializes tabular data serialized by MetaEncoder

### *class* recode.base.MetaEncoder(frame_to_chk, frame_to_meta)

Bases: [`Callable`](https://docs.python.org/3/library/collections.abc.html#collections.abc.Callable)[[[`Iterable`](https://docs.python.org/3/library/collections.abc.html#collections.abc.Iterable)[[`Any`](https://docs.python.org/3/library/typing.html#typing.Any) | [`Sequence`](https://docs.python.org/3/library/collections.abc.html#collections.abc.Sequence)[[`Any`](https://docs.python.org/3/library/typing.html#typing.Any)]]], [`bytes`](https://docs.python.org/3/builtins/stdtypes.html#bytes)]

Serializes tabular data (must be formatted as list of dicts)

### *class* recode.base.StructCodecSpecs(chk_format='d', n_channels=None, chk_size_bytes=None)

Bases: [`object`](https://docs.python.org/3/builtins/functions.html#object)

Enable the definition of codec specs based on format characters of the
python struct module
([https://docs.python.org/3/library/struct.html#format-characters](https://docs.python.org/3/library/struct.html#format-characters))

* **Parameters:**
  * **chk_format** ([`str`](https://docs.python.org/3/builtins/stdtypes.html#str)) – The format of a chunk, as specified by the struct module
    See [https://docs.python.org/3/library/struct.html#format-characters](https://docs.python.org/3/library/struct.html#format-characters)
  * **n_channels** ([`int`](https://docs.python.org/3/builtins/functions.html#int)) – Expected number of channels. If given, will assert that the
    number of channels expressed by the `chk_format` is indeed what is expected.
  * **chk_size_bytes** ([`int`](https://docs.python.org/3/builtins/functions.html#int)) – Expected number of bytes per chunk.
    If given, will assert that the chunk size expressed by the `chk_format` is
    indeed the one expected.

#### NOTE
All encoder/decoder (codec) specs can be expressed through the `chk_format`.
Yet, though `n_channels` and `chk_size_bytes` are both optional, it is advised to
include them in production code since they act as extra confirmation of the codec
to be used. Encoding and decoding problems can be hard to notice until much
later on downstream, and are therefore hard to debug.

To utilise recode, first define your codec specs. If your frame is only one channel,
then your format string will include two characters maximum: an optional special character to
control the byte order, size and alignment (@, =, <, >, !), and a format character to specify
the type of data being packed/unpacked. The format character should match the data type of the
samples in the frame so they are properly encoded/decoded.

This can be seen in the following example.

```pycon
>>> specs = StructCodecSpecs(chk_format='h')
>>> print(specs)
StructCodecSpecs(chk_format='h', n_channels=1, chk_size_bytes=2)
>>> encoder = ChunkedEncoder(frame_to_chk=specs.frame_to_chk)
>>> decoder = ChunkedDecoder(chk_to_frame=specs.chk_to_frame)
>>> frames = [1, 2, 3]
>>> b = encoder(frames)
>>> assert b == b'\x01\x00\x02\x00\x03\x00'
>>> decoded_frames = list(decoder(b))
>>> assert decoded_frames == frames
```

The only reason (but it’s a good one) to specify `n_channels` is to assert them.

```pycon
>>> specs = StructCodecSpecs(chk_format='@hh', n_channels=2)
>>> print(specs)
StructCodecSpecs(chk_format='@hh', n_channels=2, chk_size_bytes=4)
>>> encoder = ChunkedEncoder(frame_to_chk=specs.frame_to_chk)
>>> decoder = ChunkedDecoder(chk_to_frame=specs.chk_to_frame)
>>> frames = [(1, 2), (3, 4), (5, 6)]
>>> b = encoder(frames)
>>> assert b == b'\x01\x00\x02\x00\x03\x00\x04\x00\x05\x00\x06\x00'
>>> decoded_frames = list(decoder(b))
>>> assert decoded_frames == frames
```

On the other hand, if each channel has a different data type, say (int, float, int),
then your format string needs a format character for each of your channels.
This can be seen in the following example, which also shows the use
of a different byte character (=).

```pycon
>>> specs = StructCodecSpecs(chk_format = '=hdh')
>>> print(specs)
StructCodecSpecs(chk_format='=hdh', n_channels=3, chk_size_bytes=12)
>>> encoder = ChunkedEncoder(frame_to_chk = specs.frame_to_chk)
>>> decoder = ChunkedDecoder(chk_to_frame=specs.chk_to_frame)
>>> frames = [(1, 2.45, 1), (3, 4.321, 3)]
>>> b = encoder(frames)
>>> assert b == b'\x01\x00\x9a\x99\x99\x99\x99\x99\x03@\x01\x00\x03\x00b\x10X9\xb4H\x11@\x03\x00'
>>> decoded_frames = list(decoder(b))
>>> assert decoded_frames == frames
```

You can also use the IterativeDecoder which will return an iterator of frames instead of the
full list of frames, similar to what struct.iter_unpack does.
IterativeDecoder can be instantiated and called in the same way as ChunkedDecoder.
An example of IterativeDecorator can be seen below.

```pycon
>>> specs = StructCodecSpecs(chk_format = 'hdhd')
>>> print(specs)
StructCodecSpecs(chk_format='hdhd', n_channels=4, chk_size_bytes=32)
>>> encoder = ChunkedEncoder(frame_to_chk = specs.frame_to_chk)
>>> decoder = IterativeDecoder(chk_to_frame = specs.chk_to_frame)
>>> frames = [(1,1.1,1,1.1),(2,2.2,2,2.2),(3,3.3,3,3.3)]
>>> b = encoder(frames)
>>> iter_frames = decoder(b)
>>> assert next(iter_frames) == frames[0]
>>> next(iter_frames)
(2, 2.2, 2, 2.2)
```

Along with using recode for the kinds of data we have looked at so far,
it can also be applied to DataFrames when
they have been converted to a list of dicts using MetaEncoder and MetaDecoder.
An example of this can be seen below.

```pycon
>>> data = [{'foo': 1.1, 'bar': 2.2},
...         {'foo': 513.23, 'bar': 456.1},
...         {'foo': 32.0, 'bar': 6.7}]
>>> specs = StructCodecSpecs(chk_format='dd')
>>> print(specs)
StructCodecSpecs(chk_format='dd', n_channels=2, chk_size_bytes=16)
>>> encoder = MetaEncoder(frame_to_chk = specs.frame_to_chk, frame_to_meta = frame_to_meta)
>>> decoder = MetaDecoder(chk_to_frame = specs.chk_to_frame, meta_to_frame = meta_to_frame)
>>> b = encoder(data)
>>> assert decoder(b) == data
```

### *class* recode.base.codec_tuple(encode, decode)

Bases: [`tuple`](https://docs.python.org/3/builtins/stdtypes.html#tuple)

#### decode

Alias for field number 1

#### encode

Alias for field number 0

### recode.base.frame_to_meta(frame)

Defines header for serialization of tabluar data

```pycon
>>> rows = [{'customer': 1}, {'customer': 2}, {'customer': 3}]
>>> assert frame_to_meta(rows) == b'\x08\x00customer'
```

### recode.base.meta_to_frame(meta)

Deserializes header for deserialization of tabular data

```pycon
>>> meta = b'\x1c\x00customer.apple.banana.tomato\x01\x00\x01\x00\x02\x00\x03\x00\x02\x00\
... x03\x00\x02\x00\x05\x00\x01\x00\x03\x00\x04\x00\t\x00'
>>> assert meta_to_frame(meta)[0] == ['customer', 'apple', 'banana', 'tomato']
```

### recode.base.mk_codec(chk_format='d', n_channels=None, chk_size_bytes=None)

Enable the definition of codec specs based on format characters of the
python struct module
([https://docs.python.org/3/library/struct.html#format-characters](https://docs.python.org/3/library/struct.html#format-characters))

* **Parameters:**
  * **chk_format** ([`str`](https://docs.python.org/3/builtins/stdtypes.html#str)) – The format of a chunk, as specified by the struct module
    See [https://docs.python.org/3/library/struct.html#format-characters](https://docs.python.org/3/library/struct.html#format-characters)
  * **n_channels** ([`int`](https://docs.python.org/3/builtins/functions.html#int)) – Expected number of channels. If given, will assert that the
    number of channels expressed by the `chk_format` is indeed what is expected.
  * **chk_size_bytes** ([`int`](https://docs.python.org/3/builtins/functions.html#int)) – Expected number of bytes per chunk.
    If given, will assert that the chunk size expressed by the `chk_format` is
    indeed the one expected.
* **Returns:**
  A (named)tuple with encode and decode functions

```pycon
>>> from recode import mk_codec
>>> encoder, decoder = mk_codec()
>>> b = encoder([0, -3, 3.14])
>>> b
b'\x00\x00\x00\x00\x00\x00\x00\x00\x00\x00\x00\x00\x00\x00\x08\xc0\x1f\x85\xebQ\xb8\x1e\t@'
>>> decoder(b)
[0.0, -3.0, 3.14]
```

What about those channels?
Well, some times you need to encode/decode multi-channel streams, such as:

```pycon
>>> multi_channel_stream = [[3, -1], [4, -1], [5, -9]]
```

Say, for example, if you were dealing with stereo waveform
(with the standard PCM_16 format), you’d do it this way:

```pycon
>>> encoder, decoder = mk_codec('hh')
>>> pcm_bytes = encoder(iter(multi_channel_stream))
>>> pcm_bytes
b'\x03\x00\xff\xff\x04\x00\xff\xff\x05\x00\xf7\xff'
>>> decoder(pcm_bytes)
[(3, -1), (4, -1), (5, -9)]
```

The `n_channels` and `chk_size_bytes` arguments are there if you want to assert
that your number of channels and chunk size are what you expect.
Again, these are just for verification, because we know how easy it is to
misspecify the `chk_format`, and how hard it can be to notice that we did.

It is advised to use these in any production code, for the sanity of everyone!

```pycon
>>> mk_codec('hhh', n_channels=2)
Traceback (most recent call last):
  ...
AssertionError: You said there'd be 2 channels, but I inferred 3
>>> mk_codec('hhh', chk_size_bytes=3)
Traceback (most recent call last):
  ...
AssertionError: The given chk_size_bytes 3 did not match the inferred (from chk_format) 6
```

Finally, so far we’ve done it this way:

```pycon
>>> encoder, decoder = mk_codec('hHifd')
```

But see that what’s actually returned is a NAMED tuple, which means that you can
can also get one object that will have `.encode` and `.decode` properties:

```pycon
>>> codec = mk_codec('hHifd')
>>> to_encode = [[1, 2, 3, 4, 5], [6, 7, 8, 9, 10]]
>>> encoded = codec.encode(to_encode)
>>> decoded = codec.decode(encoded)
>>> decoded
[(1, 2, 3, 4.0, 5.0), (6, 7, 8, 9.0, 10.0)]
```

And you can checkout the properties of your encoder and decoder (they
should be the same)

```pycon
>>> codec.encode.chk_format
'hHifd'
>>> codec.encode.n_channels
5
>>> codec.encode.chk_size_bytes
24
```

### recode.base.mk_encoder_and_decoder(chk_format='d', n_channels=None, chk_size_bytes=None)

Enable the definition of codec specs based on format characters of the
python struct module
([https://docs.python.org/3/library/struct.html#format-characters](https://docs.python.org/3/library/struct.html#format-characters))

* **Parameters:**
  * **chk_format** ([`str`](https://docs.python.org/3/builtins/stdtypes.html#str)) – The format of a chunk, as specified by the struct module
    See [https://docs.python.org/3/library/struct.html#format-characters](https://docs.python.org/3/library/struct.html#format-characters)
  * **n_channels** ([`int`](https://docs.python.org/3/builtins/functions.html#int)) – Expected number of channels. If given, will assert that the
    number of channels expressed by the `chk_format` is indeed what is expected.
  * **chk_size_bytes** ([`int`](https://docs.python.org/3/builtins/functions.html#int)) – Expected number of bytes per chunk.
    If given, will assert that the chunk size expressed by the `chk_format` is
    indeed the one expected.
* **Returns:**
  A (named)tuple with encode and decode functions

```pycon
>>> from recode import mk_codec
>>> encoder, decoder = mk_codec()
>>> b = encoder([0, -3, 3.14])
>>> b
b'\x00\x00\x00\x00\x00\x00\x00\x00\x00\x00\x00\x00\x00\x00\x08\xc0\x1f\x85\xebQ\xb8\x1e\t@'
>>> decoder(b)
[0.0, -3.0, 3.14]
```

What about those channels?
Well, some times you need to encode/decode multi-channel streams, such as:

```pycon
>>> multi_channel_stream = [[3, -1], [4, -1], [5, -9]]
```

Say, for example, if you were dealing with stereo waveform
(with the standard PCM_16 format), you’d do it this way:

```pycon
>>> encoder, decoder = mk_codec('hh')
>>> pcm_bytes = encoder(iter(multi_channel_stream))
>>> pcm_bytes
b'\x03\x00\xff\xff\x04\x00\xff\xff\x05\x00\xf7\xff'
>>> decoder(pcm_bytes)
[(3, -1), (4, -1), (5, -9)]
```

The `n_channels` and `chk_size_bytes` arguments are there if you want to assert
that your number of channels and chunk size are what you expect.
Again, these are just for verification, because we know how easy it is to
misspecify the `chk_format`, and how hard it can be to notice that we did.

It is advised to use these in any production code, for the sanity of everyone!

```pycon
>>> mk_codec('hhh', n_channels=2)
Traceback (most recent call last):
  ...
AssertionError: You said there'd be 2 channels, but I inferred 3
>>> mk_codec('hhh', chk_size_bytes=3)
Traceback (most recent call last):
  ...
AssertionError: The given chk_size_bytes 3 did not match the inferred (from chk_format) 6
```

Finally, so far we’ve done it this way:

```pycon
>>> encoder, decoder = mk_codec('hHifd')
```

But see that what’s actually returned is a NAMED tuple, which means that you can
can also get one object that will have `.encode` and `.decode` properties:

```pycon
>>> codec = mk_codec('hHifd')
>>> to_encode = [[1, 2, 3, 4, 5], [6, 7, 8, 9, 10]]
>>> encoded = codec.encode(to_encode)
>>> decoded = codec.decode(encoded)
>>> decoded
[(1, 2, 3, 4.0, 5.0), (6, 7, 8, 9.0, 10.0)]
```

And you can checkout the properties of your encoder and decoder (they
should be the same)

```pycon
>>> codec.encode.chk_format
'hHifd'
>>> codec.encode.n_channels
5
>>> codec.encode.chk_size_bytes
24
```

### recode.base.specs_from_frames(frames)

Implicitly defines the codec specs based on the frames to encode/decode.
specs_from_frames returns a tuple of an iterator of frames and the defined StructCodecSpecs. If
frames is an iterable, then the iterator can be ignored like the following example.

```pycon
>>> frames = [1,2,3]
>>> _, specs = specs_from_frames(frames)
>>> print(specs)
StructCodecSpecs(chk_format='h', n_channels=1, chk_size_bytes=2)
>>> encoder = ChunkedEncoder(frame_to_chk = specs.frame_to_chk)
>>> decoder = ChunkedDecoder(chk_to_frame=specs.chk_to_frame)
>>> b = encoder(frames)
>>> assert b == b'\x01\x00\x02\x00\x03\x00'
>>> decoded_frames = list(decoder(b))
>>> assert decoded_frames == frames
```

If frames is an iterator, then we can still use specs_from_frames as long as we redefine frames
from the output like in the following example.

```pycon
>>> frames = iter([[1.1,2.2],[3.3,4.4]])
>>> frames, specs = specs_from_frames(frames)
>>> print(specs)
StructCodecSpecs(chk_format='dd', n_channels=2, chk_size_bytes=16)
>>> encoder = ChunkedEncoder(frame_to_chk = specs.frame_to_chk)
>>> decoder = ChunkedDecoder(chk_to_frame=specs.chk_to_frame)
>>> b = encoder(frames)
>>> decoded_frames = list(decoder(b))
>>> assert decoded_frames == [(1.1,2.2),(3.3,4.4)]
```


# _autosummary/recode.html.md

# recode

Make codecs for fixed size structured chunks serialization and deserialization of
sequences, tabular data, and time-series.

The easiest and bigest bang for your buck is `mk_codec`

```pycon
>>> from recode import mk_codec
>>> encoder, decoder = mk_codec()
```

`encoder` will encode a list (or any iterable) of numbers into bytes

```pycon
>>> b = encoder([0, -3, 3.14])
>>> b
b'\x00\x00\x00\x00\x00\x00\x00\x00\x00\x00\x00\x00\x00\x00\x08\xc0\x1f\x85\xebQ\xb8\x1e\t@'
```

`decoder` will decode those bytes to get you back your numbers

```pycon
>>> decoder(b)
[0.0, -3.0, 3.14]
```

There’s only really one argument you need to know about in `mk_codec`.
The first argument, called `chk_format`, which is a string of characters from
the “Format” column of
[https://docs.python.org/3/library/struct.html#format-characters](https://docs.python.org/3/library/struct.html#format-characters)

The length of the string specifies the number of “channels”,
and each individual character of the string specifies the kind of encoding you should
apply to each “channel” (hold your horses, we’ll explain).

The one we’ve just been through is in fact

```pycon
>>> encoder, decoder = mk_codec('d')
```

That is, it will expect that your data is a list of numbers, and they’ll be encoded
with the ‘d’ format character, that is 8-bytes doubles.
That default is goo because it gives you a lot of room, but if you knew that you
would only be dealing with 2-byte integers (as in most WAV audio waveforms),
you would have chosen `h`:

```pycon
>>> encoder, decoder = mk_codec('h')
```

What about those channels?
Well, some times you need to encode/decode multi-channel streams, such as:

```pycon
>>> multi_channel_stream = [[3, -1], [4, -1], [5, -9]]
```

Say, for example, if you were dealing with stereo waveform
(with the standard PCM_16 format), you’d do it this way:

```pycon
>>> encoder, decoder = mk_codec('hh')
>>> pcm_bytes = encoder(iter(multi_channel_stream))
>>> pcm_bytes
b'\x03\x00\xff\xff\x04\x00\xff\xff\x05\x00\xf7\xff'
>>> decoder(pcm_bytes)
[(3, -1), (4, -1), (5, -9)]
```

The `n_channels` and `chk_size_bytes` arguments are there if you want to assert
that your number of channels and chunk size are what you expect.
Again, these are just for verification, because we know how easy it is to
misspecify the `chk_format`, and how hard it can be to notice that we did.

It is advised to use these in any production code, for the sanity of everyone!

```pycon
>>> mk_codec('hhh', n_channels=2)
Traceback (most recent call last):
  ...
AssertionError: You said there'd be 2 channels, but I inferred 3
>>> mk_codec('hhh', chk_size_bytes=3)
Traceback (most recent call last):
  ...
AssertionError: The given chk_size_bytes 3 did not match the inferred (from chk_format) 6
```

Finally, so far we’ve done it this way:

```pycon
>>> encoder, decoder = mk_codec('hHifd')
```

But see that what’s actually returned is a NAMED tuple, which means that you can
can also get one object that will have `.encode` and `.decode` properties:

```pycon
>>> codec = mk_codec('hHifd')
>>> to_encode = [[1, 2, 3, 4, 5], [6, 7, 8, 9, 10]]
>>> encoded = codec.encode(to_encode)
>>> decoded = codec.decode(encoded)
>>> decoded
[(1, 2, 3, 4.0, 5.0), (6, 7, 8, 9.0, 10.0)]
```

And you can checkout the properties of your encoder and decoder (they
should be the same)

```pycon
>>> codec.encode.chk_format
'hHifd'
>>> codec.encode.n_channels
5
>>> codec.encode.chk_size_bytes
24
```

### Modules

| [`audio`](_autosummary/recode.audio.html.md#module-recode.audio)   | Encoding audio                       |
|------------------------------------------------------------------------------|--------------------------------------|
| [`base`](_autosummary/recode.base.html.md#module-recode.base)     | Base recode objects                  |
| [`util`](_autosummary/recode.util.html.md#module-recode.util)     | Utils for use throughout the package |


# _autosummary/recode.util.html.md

# recode.util

Utils for use throughout the package

### Functions

| [`get_struct`](_autosummary/recode.util.html.md#recode.util.get_struct)(str_type)      |    |
|----------------------------------------------------------------------------|----|
| [`list_of_dicts`](_autosummary/recode.util.html.md#recode.util.list_of_dicts)(cols, vals) |    |
| [`spy`](_autosummary/recode.util.html.md#recode.util.spy)(iterable[, n])        |    |
| [`take`](_autosummary/recode.util.html.md#recode.util.take)(n, iterable)         |    |

### recode.util.get_struct(str_type)

```pycon
>>> assert get_struct(type(1)) == 'h'
>>> assert get_struct(type(1.001)) == 'd'
```

### recode.util.list_of_dicts(cols, vals)

```pycon
>>> cols = ['foo', 'bar']
>>> vals = [[1,2], [3,4], [5,6]]
>>> list_of_dicts(cols, vals)
[{'foo': 1, 'bar': 2}, {'foo': 3, 'bar': 4}, {'foo': 5, 'bar': 6}]
```

### recode.util.spy(iterable, n=1)

```pycon
>>> peek, it = spy([1,2,3], 1)
>>> assert peek == [1]
>>> assert next(it) == 1
>>> assert list(it) == [2,3]
```

### recode.util.take(n, iterable)

```pycon
>>> assert take(3, [1,2,3,4,5]) == [1,2,3]
```


# about-this-build.html.md

<!-- generated by epythet -->

# About this build

This documentation was built on **2026-09-22 14:10 UTC** from commit <a href="https://github.com/i2mint/recode/commit/85e0943c3489f32ef7efc59e10d5880ee7e312f3"><code>85e0943</code></a> on branch <code>master</code>, for **recode 0.1.40** (from <code>setup.cfg</code>).

#### NOTE
Nothing suggests a mismatch: the tree was clean at the commit above, and the documented version is the one on PyPI.

## Source

|                     |                                                                                                                                                      |
|---------------------|------------------------------------------------------------------------------------------------------------------------------------------------------|
| Commit              | <a href="https://github.com/i2mint/recode/commit/85e0943c3489f32ef7efc59e10d5880ee7e312f3"><code>85e0943c3489f32ef7efc59e10d5880ee7e312f3</code></a> |
| Branch              | <code>master</code>                                                                                                                                  |
| Tags at this commit | <code>0.1.40</code>                                                                                                                                  |
| Working tree        | clean                                                                                                                                                |
| Remote              | <code>https://github.com/i2mint/recode</code>                                                                                                        |

## Continuous integration

|              |                                                                                            |
|--------------|--------------------------------------------------------------------------------------------|
| Repository   | <code>i2mint/recode</code>                                                                 |
| Run          | <a href="https://github.com/i2mint/recode/actions/runs/35738301691">35738301691</a>        |
| Ref          | <code>refs/heads/master</code>                                                             |
| Event commit | <code>3b92016bb6dcf47b543cb464a8b6b4420be08e08</code> (in the history of the built commit) |

## Tools

|          |         |
|----------|---------|
| epythet  | 0.2.12  |
| Sphinx   | 9.1.0   |
| docutils | 0.22.4  |
| Python   | 3.12.14 |

## Configuration as resolved

|               |                                                                  |
|---------------|------------------------------------------------------------------|
| theme         | <code>auto</code> (Sphinx theme <code>furo</code>)               |
| accent        | <code>#58489b</code>                                             |
| api_generator | <code>autosummary</code>                                         |
| ignore        | <code>tests/</code>, <code>scrap/</code>, <code>examples/</code> |
| agent_outputs | <code>true</code>                                                |
| aggregates    | <code>md</code>                                                  |
| ai_artifacts  | <code>true</code>                                                |

## Package on PyPI

Latest release: <a href="https://pypi.org/project/recode/0.1.40/">0.1.40</a>, the same as the documented version.

## Reproduce

```bash
git clone https://github.com/i2mint/recode && cd recode
git checkout 85e0943c3489f32ef7efc59e10d5880ee7e312f3
pip install "epythet==0.2.12"
epythet quickstart . --ignore tests/ scrap/ examples/
```

The same data, for machines: <a href="build_info.json"><code>build_info.json</code></a> (schema version 1).


# api.html.md

# API reference

| [`recode`](_autosummary/recode.html.md#module-recode)   | Make codecs for fixed size structured chunks serialization and deserialization of sequences, tabular data, and time-series.   |
|-------------------------------------------------------------------------|-------------------------------------------------------------------------------------------------------------------------------|


