Comments (4)
Hi @mmaaz60
Thank you for your interest in MDETR.
It looks like you training diverged. Can I ask how many gpus you used?
from mdetr.
Hi @mmaaz60
Thank you for your interest in MDETR.
It looks like you training diverged. Can I ask how many gpus you used?
Thank You @alcinos,
I used 32 GPUs with batch_size
of 2 per GPU.
from mdetr.
Hum that’s quite surprising then. Nothing fishy happened, like the job getting preempted then restarted?
Are you sure you have the correct transformers version?
Otherwise mb try with a slightly smaller lr?
from mdetr.
Thank You
Hum that’s quite surprising then. Nothing fishy happened, like the job getting preempted then restarted?
Nothing such happened during training
Are you sure you have the correct transformers version?
I am using transformers version 4.5.1
Otherwise mb try with a slightly smaller lr?
I actually stopped and then resumed the training from the 19th epoch and now it reaches to 25th epoch and seems to be converging. Not sure what went wrong previously as I didn't change anything when resuming.
from mdetr.
Related Issues (20)
- Can you provide the final_vg json files?
- Issues with training on a single node
- requirements for running the demo notebook HOT 2
- Minor Typo in eval_lvis.py
- Generate tokens_positive for just COCO for pre-training!!
- 2D Object Detection of Drone Images
- Bbox assertion error when using ENB models (eval & pretrain as well): assert (boxes1[:, 2:] >= boxes1[:, :2]).all() HOT 8
- how to define positive tokens? HOT 1
- ValueError: char_to_token() is not available when using Python based tokenizers HOT 1
- How to generate "tokens_negative" and "tokens_positive" when we convert our own dataset into mdetr annotations? HOT 1
- finetune have bug!!ValueError: char_to_token() is not available when using Python based tokenizers HOT 3
- Cannot find finetune_phrasecut_miniv.json HOT 1
- [Colab Error] InvalidVersion: Invalid version: '0.10.1,<0.11' HOT 2
- With the increase of MDETR training time, GPU memory occupation keeps increasing. After several epoch training, memory explodes,i.e. OOM, out of memory. HOT 1
- issue #44 "I guess its uncompleted project so I gave up this essay" HOT 1
- I would like to know how to use the model in clevr-ref+.
- Weird results while evaluating & reproducing ENB3 model on PhraseCut
- How to decode GQA evaluation results HOT 1
- run
- CLEVR-REF+ training
Recommend Projects
-
React
A declarative, efficient, and flexible JavaScript library for building user interfaces.
-
Vue.js
🖖 Vue.js is a progressive, incrementally-adoptable JavaScript framework for building UI on the web.
-
Typescript
TypeScript is a superset of JavaScript that compiles to clean JavaScript output.
-
TensorFlow
An Open Source Machine Learning Framework for Everyone
-
Django
The Web framework for perfectionists with deadlines.
-
Laravel
A PHP framework for web artisans
-
D3
Bring data to life with SVG, Canvas and HTML. 📊📈🎉
-
Recommend Topics
-
javascript
JavaScript (JS) is a lightweight interpreted programming language with first-class functions.
-
web
Some thing interesting about web. New door for the world.
-
server
A server is a program made to process requests and deliver data to clients.
-
Machine learning
Machine learning is a way of modeling and interpreting data that allows a piece of software to respond intelligently.
-
Visualization
Some thing interesting about visualization, use data art
-
Game
Some thing interesting about game, make everyone happy.
Recommend Org
-
Facebook
We are working to build community through open source technology. NB: members must have two-factor auth.
-
Microsoft
Open source projects and samples from Microsoft.
-
Google
Google ❤️ Open Source for everyone.
-
Alibaba
Alibaba Open Source for everyone
-
D3
Data-Driven Documents codes.
-
Tencent
China tencent open source team.
from mdetr.