The algorithm is ported over from Chromium's, which produces very good
results for speech, while being not objectionable for music, especially
in the background.
Other algorithms were tested.
Phase vocoders:
- Bungee
- Signalsmith Stretch
- pvdoneright
All three of these exhibited the usual phase shifting artifacts,
causing speech to sound slightly shifted into the high end. Speech is
the main thing we want to optimize for, so these aren't ideal.
Sonic (TD-PSOLA) performs better than WSOLA for speech, especially at
rates higher than 2x, but makes background sounds/music garbled and
unpleasant. It is still worth considering for speech clarity, and could
be added as an optional feature.
This allows nodes downstream of the mixer to know the exact frame at
which EOS is reached.
More concepts have been introduced into PipelineStatus.h to make each
condition in the pipeline clearer about its intent.
This allows the downstream node to query status and sleep again
synchronously, which is needed for nodes that need multiple input
blocks before they can produce output and wake their downstream nodes.
We could potentially have a pull() that would output audio after a
status query said that data was pending. Instead, ensure that we keep
the subsequent pull() from providing anything in such cases.
Also, rewrite HaveData to Pending when it outputs nothing. When
multiple audio tracks were enabled, it was possible to reach the empty
block branch with a HaveData status, since it takes priority over
Pending. This fixes a crash in pull() verifying that the combined
status is not HaveData.
Seeks don't always move a decoded data producer's head, so we need to
make sure not to remove queued data downstream when that is the case.
To communicate this, the producers can now be queried before pulling
data, allowing them to have an in-band signal to clear the queued data
after a seek has moved the producer and broken monotonicity.
This fixes a flake in HTMLVideoElement-resize-event-during-playback.
This isn't strictly an optimization, but at this stage that is its
effect. By not signaling unless the combined state has changed, we
avoid waking the downstream node unnecessarily. However, the real
purpose is to allow us to transmit a new signal when the upstream
position moves to clear the data downstream to allow repeat timestamps.
Previously, we weren't too consistent about the definition of frame and
sample when it relates to raw audio data. This brings all the usages in
the context of raw data in line (hopefully), with samples referring to
a single PCM value, and frames referring to the multiple samples that
make up an instant's audio across all channels.
Now, all nodes are connected through Sink::connect_input() and
disconnect_input().
AudioMixer now derives from a base AudioProcessor class that inherits
from both AudioSink and AudioProducer. It is the only current node that
can accept multiple inputs, tracking each one by its pointer identity.
Seeking is now unified under one single method signature implemented by
all producers and transmitted through the pipeline by all sinks. By
doing it this way, we can simply instantaneously notify each node of
the pipeline that it needs to stop what it's doing and seek. For nodes
that are threaded (particularly the source providers), this causes them
to stop pushing data to their queue immediately, so that no stale data
makes it through to the output. Then, when new data does come through,
that is a clear indication that the seek has completed.
Note that track enablement is now through the pipeline as well, which
means that SuspendedStateHandler no longer has a way to suspend newly-
enabled tracks. Decoder suspension will need to be reworked to fit into
this new pipeline, sleeping/disposing and restarting entirely based on
the pull() timing in the producers.
Using a callback shared by all producers in the pipeline, notify the
AudioPlaybackSink when it needs to wake up and start processing data
again. Prior to this commit, it was simply burning CPU spinning until
data was produced.
This is an intermediate step towards unifying the pipeline around new
Producer/Sink interfaces. Producers now have a pull() method that gets
the next piece of data from them. The pull() method returns a status
that can indicate whether it has current data, and if not, why it's
unavailable. This signal will be passed down the pipeline to the final
sink, which can expose the signal to its user, which in the normal
playback pipeline is PlaybackManager. The signal can be used to
transition between playback states. Currently, this is only hooked up
to the buffering state, but should be used later for ending playback
as well as decoding error propagation.
Buffering is now determined solely based on whether the pipeline is
blocked on incomplete data, so the ready state for video now progresses
past HAVE_METADATA immediately after playback manager initializes. This
will change when files have buffered ranges.
This will allow reuse of the allocations, instead of reallocating a
FixedArray for every block that changes size. Generally, it will be
possible to reuse AudioBlock memory throughout most of the pipeline
at least while in a single process.
- Provider -> producer
- (Audio|Video)DataProvider -> Decoded(Audio|Video)Producer
- MediaTimeProvider remains suffixed Provider, moves out of the
Providers folder to the root of LibMedia
This brings the naming more in line with the intended split
functionality split between different nodes in the pipeline.
This doesn't actually change things too much from the prior commit, but
acts as a step towards making mixing into a sink/provider combo in the
new pipeline model.