Comments (5)
That Nemo version is 6 months old, can you use r1.23 and see if it persists ? We do not see constantly increasing CPU memory per epoch, but that may be because we use multiple nodes - min 4 nodes
from nemo.
Is it GPU or CPU memory that is exhausted ? And how many nodes are you using ?
What version of NeMo are you using ?
Without sufficient details it's not possible to debug.
What I can say is we train on nodes with 400 GB ram per node and A100 with 80GB gpu memory and train on 90-400K hours of speech without oom in either CPU or GPU memory.
If you can visibly see CPU ram constantly increase during training, a pseudo fix could be to use exp_manager.max_time_per_run and set it to a reasonable value like a day, then the job stops after a day and you can restart it and avoid memory leak. It's not a fix but a temporary solution
from nemo.
Is it GPU or CPU memory that is exhausted ? And how many nodes are you using ?
- CPU not GPU
- Just single node
We just added one row
self.log('loss', loss_value, on_step=True, prog_bar=True, on_epoch=False,)
in file:
nemo/collections/asr/models/ctc_models.py
Previously, we used on_epoch=True, but now the problem still remains after changine to False.
What version of NeMo are you using ?
git log
commit 0d3d8fa (HEAD -> main)
Author: anteju [email protected]
Date: Wed Nov 15 16:56:29 2023 -0800
[ASR] GSS-based mask estimator (#7849)
* Added GSS-based mask estimator for multispeaker scenarios
Signed-off-by: Ante Jukić <[email protected]>
* Addressed PR comments
Signed-off-by: Ante Jukić <[email protected]>
---------
Signed-off-by: Ante Jukić <[email protected]>
Co-authored-by: Taejin Park <[email protected]>
Actually, it's very easy to verify: you just submit a training task with, say librispeech data, you can observe you CPU memory keeps increasing within an epoch.
But such memory increase won't hurt since memory increase slow and after an epoch, memory usage somehwo is going down again. Here, if we decrease our training data down to 30k, for 1.2T cpu memory, we can finish an epoch normally.
from nemo.
from nemo.
Hi, is this issue resolved? I've been running into the same issue. (I can confirm that it happens on 1.23 as well)
from nemo.
Related Issues (20)
- Not found file "convert_mistral_hf_to_nemo.py" in /opt/NeMo/scripts/checkpoint_converters/ for Convert Mistral HOT 1
- Precision Problem between nemo model and hugging face model HOT 2
- Llama2 70B SFT with FSDP failing HOT 2
- training config used for training stt_en_quartznet15x5 HOT 2
- llama2 training hangs when pp_size > 1 HOT 2
- Integration of Turn-Taking Models into Nemo Framework for Enhanced Realistic Conversations
- FileNotFoundError: Model stt_fa_fastconformer_hybrid_large was not found. HOT 6
- [Feature] Add Support on Multiple Metrics Reporting during Training Progress for Validation
- checkpoints not saved due to wrong loss comparison?
- when "write_predictions_to_file" is true,generate will fail。 HOT 3
- "RuntimeError: start (4) + length (1) exceeds dimension size (4)." when running cache aware streaming inference
- slow validation process HOT 2
- Optimizing Learning Rate Parameters in Model Fine-tuning HOT 1
- AUDIO FILE SIZE for fine tuning STT En FastConformer Hybrid Transducer-CTC Large Streaming Multi HOT 1
- `EncDecCTCModel.transcribe(audio=...)` changed to `EncDecCTCModel.transcribe(paths2audio_files=...)` HOT 7
- Enormous number of `.nemo` checkpoints produced in training HOT 4
- [Conversion] How to convert Finetuned T5 checkpoint ended with `.ckpt` to `.nemo` checkpoint with NeMo toolkit?
- Can't launch NeMo containers with CUDA support
- Latest huggingface transformers version breaking nlp modules HOT 6
- Any tts models in nemo that can simulated human laughter and other human sounds?
Recommend Projects
-
React
A declarative, efficient, and flexible JavaScript library for building user interfaces.
-
Vue.js
🖖 Vue.js is a progressive, incrementally-adoptable JavaScript framework for building UI on the web.
-
Typescript
TypeScript is a superset of JavaScript that compiles to clean JavaScript output.
-
TensorFlow
An Open Source Machine Learning Framework for Everyone
-
Django
The Web framework for perfectionists with deadlines.
-
Laravel
A PHP framework for web artisans
-
D3
Bring data to life with SVG, Canvas and HTML. 📊📈🎉
-
Recommend Topics
-
javascript
JavaScript (JS) is a lightweight interpreted programming language with first-class functions.
-
web
Some thing interesting about web. New door for the world.
-
server
A server is a program made to process requests and deliver data to clients.
-
Machine learning
Machine learning is a way of modeling and interpreting data that allows a piece of software to respond intelligently.
-
Visualization
Some thing interesting about visualization, use data art
-
Game
Some thing interesting about game, make everyone happy.
Recommend Org
-
Facebook
We are working to build community through open source technology. NB: members must have two-factor auth.
-
Microsoft
Open source projects and samples from Microsoft.
-
Google
Google ❤️ Open Source for everyone.
-
Alibaba
Alibaba Open Source for everyone
-
D3
Data-Driven Documents codes.
-
Tencent
China tencent open source team.
from nemo.