Respiratory inference explorer
Watch a 1D convolution scan a waveform, inspect the matrix arithmetic and feature maps, then follow the 168-feature Edge TPU pipeline through six development phases.
Machine learning · Edge deployment
Completed research prototypeA six-phase research project that advanced respiratory-audio classification from data preparation and CNN experiments to real-time Google Coral Edge TPU inference.
Interactive project overview
These demonstrations present the project's state, timing, and key tradeoffs. The full case study and implementation details follow below.
Watch a 1D convolution scan a waveform, inspect the matrix arithmetic and feature maps, then follow the 168-feature Edge TPU pipeline through six development phases.
The team set out to classify irregular respiratory sounds locally so audio capture, preprocessing, inference, and result reporting would not depend on a cloud service.
Each target—from a 264 KB microcontroller to an ESP32 and the Coral Dev Board—changed the viable input representation, runtime, memory budget, conversion path, and peripheral strategy.
Inspect the final audio pipeline and the hardware constraints that drove changes in model representation, deployment tooling, and device selection.
Interactive architecture
The architecture changed as data quality, RAM, peripheral integration, and Edge TPU compatibility exposed new constraints.
Selected component
Annotations were parsed into normal, crackle, wheeze, and combined clips; later filtering retained higher-quality samples.
The components I designed, implemented, tested, or integrated.
Built Python preprocessing pipelines for waveform slicing, spectrograms, Mel spectrograms, MFCCs, filtering, and dataset quality control.
Created and optimized TensorFlow/Keras and PyTorch CNNs through optimizer experiments, LeakyReLU and SpatialDropout revisions, 1D convolution, temporal attention, and audio augmentation.
Quantized models for constrained hardware and diagnosed why the first respiratory model fit flash but exceeded the Raspberry Pi Pico's runtime RAM budget.
Built the PC-to-Pico UART inference path and the ESP32 microphone, LCD, Wi-Fi, Flask, and on-device inference prototype.
Retrained the final model in a compatible TensorFlow environment, completed full-integer conversion, compiled it for the Edge TPU, and integrated live Coral inference.
The constraints and tradeoffs that shaped the implementation.
Constraint. The project needed an audio representation that made useful patterns learnable without assuming the first feature type was best.
Decision. Compare spectrogram, Mel spectrogram, and MFCC pipelines. Mel spectrograms produced the strongest initial speech result and the best respiratory result with Adam.
Constraint. Generating and processing a 2D representation on a constrained target increased memory and preprocessing cost.
Decision. Develop a 1D convolution path over raw waveforms, then add temporal attention and controlled augmentation to improve focus and robustness.
Constraint. The quantized respiratory model could be stored on the Pico but could not coexist in 264 KB of RAM with the interpreter and tensors.
Decision. Validate the microcontroller pipeline with the smaller speech model and move the final respiratory workload to hardware with sufficient memory.
Constraint. The final public model's TensorFlow version was incompatible with the Coral compiler.
Decision. Create a Python/TensorFlow 2.5 environment, retrain and quantize the network, then compile a full-integer model for the Edge TPU delegate.
Important revisions, technical pivots, and lessons from each stage.
Phase 01
Built the feature pipeline and an eight-keyword CNN, reached about 90% validation accuracy with Mel spectrograms, reduced the model from 18.6 MB to 1.55 MB, and demonstrated live inference.
Phase 02
Parsed annotated respiratory audio, tested feature/dataset/optimizer combinations, filtered to 44.1 kHz samples, and pushed the revised LeakyReLU/SpatialDropout model above the required 70% threshold.
Phase 03
Diagnosed the respiratory model's RAM failure, pivoted to the smaller speech model, and completed a PC-to-Pico UART pipeline using TensorFlow Lite for Microcontrollers.
Phase 04
Removed on-device spectrogram generation, introduced multi-head temporal attention with a residual connection, and added noise, gain, shift, and polarity transformations.
Phase 05
Validated microphone, LCD, and Wi-Fi peripherals separately, then combined recording, inference, display, and server reporting into a button-driven embedded workflow.
Phase 06
Tested a printed stethoscope, corrected data and compatibility issues, rebuilt the model for Edge TPU compilation, and delivered a self-contained six-class inference loop.
Measurements, configuration boundaries, and outcomes that show the scope of the work.
Initial validation
≈90%
Validation accuracy of the eight-keyword speech-command CNN using Mel spectrograms.
Quantized model
18.6 → 1.55 MB
Post-training quantization reduced the speech-command model by about 91.7% with minimal reported accuracy loss.
Respiratory threshold
>70%
The revised respiratory model exceeded the required accuracy target after dataset and architecture changes.
Pico RAM boundary
264 KB
The model fit flash, but the interpreter and runtime tensors could not fit the target's available RAM.
Final feature vector
168
Forty MFCC means plus 128 Mel-band means per four-second, 16 kHz capture.
Final output
6 classes
COPD, healthy, URTI, bronchiectasis, pneumonia, and bronchiolitis probabilities.
Focused excerpts paired with the engineering behavior each one implements.
The script evaluated seven TensorFlow optimizers across multiple audio representations and dataset formats.
01optimizers = {02 'Adam': tf.keras.optimizers.Adam(),03 'SGD': tf.keras.optimizers.SGD(),04 'RMSprop': tf.keras.optimizers.RMSprop(),05 'Adagrad': tf.keras.optimizers.Adagrad(),06 'Adadelta': tf.keras.optimizers.Adadelta(),07 'Adamax': tf.keras.optimizers.Adamax(),08 'Nadam': tf.keras.optimizers.Nadam()09}10 11for optimizer in optimizers.values():12 model.compile(13 optimizer=optimizer,14 loss='categorical_crossentropy',15 metrics=['accuracy']16 )The 1D architecture reshapes convolution features into a temporal sequence, applies multi-head attention, and adds the attended representation through a residual connection.
01def forward(self, x):02 x = self.conv_layers(x)03 04 # (batch, channels, time) -> (batch, time, channels)05 x = x.permute(0, 2, 1)06 07 attention_out = self.temporal_attention(x)08 x = x + attention_out09 10 # (batch, time, channels) -> (batch, channels, time)11 x = x.permute(0, 2, 1)12 x = self.classifier(x)13 14 return xThe training pipeline randomly selects noise, time shift, gain, or polarity inversion while keeping the source sample and label unchanged.
01train_transform = RandomChoice(02 transforms=[03 AddGaussianNoise(snr_db=20),04 TimeShift(05 max_shift_ms=100,06 sample_rate=sample_rate07 ),08 RandomGain(),09 PolarityInversion()10 ]11)12 13class AugmentedDataset(Dataset):14 def __getitem__(self, idx):15 original_idx = idx // (1 + self.num_augmented)16 copy_type = idx % (1 + self.num_augmented)17 waveform, label = self.base_dataset[original_idx]18 19 if copy_type > 0 and self.transform:20 waveform = self.transform(waveform)21 22 return waveform, labelThe final Coral pipeline concatenates the mean of 40 MFCC channels with 128 Mel bands, checks the expected size, and reshapes the vector for inference.
01def extract_features_from_data(02 audio_data,03 sample_rate,04 n_mfcc=40,05 n_mels=12806):07 mfccs = librosa.feature.mfcc(08 y=audio_data,09 sr=sample_rate,10 n_mfcc=n_mfcc11 )12 mel = librosa.feature.melspectrogram(13 y=audio_data,14 sr=sample_rate,15 n_mels=n_mels16 )17 18 mfccs_mean = np.mean(mfccs, axis=1)19 mel_mean = np.mean(mel, axis=1)20 features = np.concatenate((mfccs_mean, mel_mean))21 22 return featuresThe final loop records a four-second clip, filters it, extracts features, invokes the compiled TFLite model, and normalizes six disease-class scores.
01filtered_audio = butter_bandpass_filter(02 audio_chunk,03 LOWCUT,04 HIGHCUT,05 SAMPLE_RATE,06 order=FILTER_ORDER07)08 09features = extract_features_from_data(10 filtered_audio,11 SAMPLE_RATE12)13 14input_data = features.reshape(15 1,16 INPUT_FEATURE_SIZE,17 118).astype(np.float32)19 20interpreter.set_tensor(21 input_details[0]['index'],22 input_data23)24interpreter.invoke()25 26output_data = interpreter.get_tensor(27 output_details[0]['index']28)[0]29probabilities = softmax(output_data)Public links open in a new tab. Private code and project artifacts are summarized without exposing infrastructure details or credentials.