Randomly stopping /...
 
Share:
Notifications
Clear all

UPDATE FROM MABULA

I have had a very rough 2 months unfortunately health wise. I was struck 3x in a row with bacterial infections. The second infection occurred  after a routine hospital checkup and I became very sick. I had to rest and take a lot of antibiotics. Once I was recovering and restarting work 1 month ago, I again became sick. The infection was not yet gone, so even more antibiotics and rest was needed. Needless to say, it took a lot of my energy and I needed a lot of rest.

 

Finally, the infection is really gone and my energy is coming back step-by-step now and I have started work again. I am terribly sorry to have kept you waiting for support. I will address all outstanding questions and e-mails step-by step and will be on the forum daily from today.

 

APP 2.0.0-beta47 will come soon as well, a lot of work was already completed before I became very ill, so the release is also nearly ready for you. Beta47 will be much faster actually. Many workflows will be more than 2x faster, mosaics can even be 10x faster than before because registration really received a major boost... all compared to beta46.  Before I release it, I will make sure that everything is working properly and then I will release it.

Randomly stopping / freezing during integrations

17 Posts
2 Users
0 Reactions
590 Views
(@wheeljack)
Red Giant
Joined: 7 years ago
Posts: 40
Topic starter  

Not really sure how to best explain this, but I notice that frequently, the program just... halts randomly in the middle of integration. I am now trying to stack 506 images for the third time. First time it stopped at 186/506, now it stopped processing at 104/506. 

Sometimes I can still interact with the menu and the "cancel" button will works so the program itself hasn't frozen, but it just stops processing. Other times, pressing cancel does nothing and I have to exit the entire program. By doing so, I have to start over adding all of the sessions again. 

This time, pressing cancel (seemingly) did nothing, but I kept it in the background while typing this post, heard the "GONG" in the background and the process had cancelled. Took several minutes from pressing cancel until it actually cancelled. 

image

 

Sometimes it goes through fine, usually during the second attempt, but this time I am now at my third. I experienced this with the previous version too (Beta39), but only "now and then". 

Running Widows 11 25H2 ver 8246

64 GB RAM 

AMD Ryzen 9 7950x 16-core 

Usually it integrates 1, maybe 2 frames per second with that setup, until it suddenly stops dead in its tracks and goes several minutes without doing anything. Last time I waited for 10 minutes to see if a background process was catching up, but it was left at 186/506 for all that time. 

Anyone else experienced similar and know of a fix? 



   
ReplyQuote
(@wheeljack)
Red Giant
Joined: 7 years ago
Posts: 40
Topic starter  

Third time it went through, but now it has instead been stuck here for 20 minutes at 100% CPU with no apparent progress

image

Will let it run for 5 more minutes then I cancel, because I am fearing for the health of my CPU at this point. 

I have not had any problems like this earlier, but I don't know how to troubleshoot this either. 

 



   
ReplyQuote
(@mabula-admin)
Universe Admin
Joined: 9 years ago
Posts: 5369
 

Hi Wheeljack, @wheeljack

Thank you very much for reporting this. That all sounds a bit like a problem with internal memory management getting stuck/needing to do a lot of cleaning work but not able to do it somehow. I will have a good look and will do some related testing. We have moved to a newer development platform and that would explain this somehow. I can send you a test version that will do this differently okay? Then you can test if it works as expected. 

Only running APP on Windows? (then I only need to make a windows test version for you)

Mabula

 



   
ReplyQuote
(@wheeljack)
Red Giant
Joined: 7 years ago
Posts: 40
Topic starter  

Wow! That was quick. 

Yes, only using Windows. I can happily test for you if needed. 



   
ReplyQuote
(@mabula-admin)
Universe Admin
Joined: 9 years ago
Posts: 5369
 

Great, thanks @wheeljack,

I will try to provide a test version in the coming days for you. I will change the internal memory configuration and then we should be able to make it work stable again.

Mabula



   
ReplyQuote
(@wheeljack)
Red Giant
Joined: 7 years ago
Posts: 40
Topic starter  

Update: I saw that beta42 was available, so I went ahead and tried that one. 

This is indeed major improvements, it seemed to (on average) use less resources but not at the cost of speed. I went all-in and added 724 light-frames (180s each), from 7 different sessions, each with their own calibration frames. I also used 4th degree LNC with 3 iterations. 

The process started fine, but ran into issues near the end. 

The integration process stopped briefly for almost 30 seconds around frame 530-ish, but eventually continued and then it froze for almost 2 minutes at roughly 620 frames before it again froze for a minute at frame 626

image

When it continued, it again halted at 644 frames and was stuck there for 3 minutes 

image

Went at the normal speed again (roughly 1-2 frames pr second) and again it froze up at 661 frames for another minute and a half.

image

After that, it froze for about 1-3 minutes every 10-ish frames until it completed the integration. 

It went through the first LNC iteration, but has now been stuck at the calculation for 15 minutes with no apparent progress. 

image

 

It appears to no longer be trying to do something, as my CPU fans calmed down and APP is hardly using any resources anymore. Yet  It still hasn't progressed any further so I manually cancelled after about an hour and a half (in total)


This post was modified 4 months ago 2 times by Wheeljack

   
ReplyQuote
(@mabula-admin)
Universe Admin
Joined: 9 years ago
Posts: 5369
 

Hi @wheeljack,

Thank you very much for sharing this. 

Hmm, that really sounds like a big memory issue internally somehow, think the internally memory gets very fragmented. Can you possible upload the whole dataset so I can give it a spin on different computers with different Operating Systems.

I can also change some internal memory settings and provide you a test version tomorrow to see if that solves it?

What do you prefer? If you want to upload the data, I will work on it as soon as it is uploaded.

https://upload.astropixelprocessor.com/

username: uploadData

password: uploadTestData

Please make a folder with your name and issue like : Wheeljack-process-stalls

and upload the data there.

Thanks!

Mabula

 

 



   
ReplyQuote
(@wheeljack)
Red Giant
Joined: 7 years ago
Posts: 40
Topic starter  

Thanks alot Mabula. 

They're on their way now, but I reached a upload-limit so I couldn't upload any of the darks and only half of the Bias frames. (Not sure if they're even needed for your testing though so they may not even be that important) 



   
ReplyQuote
(@mabula-admin)
Universe Admin
Joined: 9 years ago
Posts: 5369
 

Hi Wheeljack,

Great, I already did some testing today on a set with 500 images. I think I know what the issue is. The recent optimizations have made APP quite a bit faster and now memory clean up becomes a big problem with the current internal memory configuration. So the clean up can not hold on to the amount of GBs that are processed. The current configuration was indeed completely fragmenting internal memory and the stalls then start to appear because of the fragmentation.

This sounds dramatic, but I think that I have already good news 🙂

There is a way to avoid fragmentation completely with a different internal approach. I was just testing this afternoon and it seems to be much better already. I was able to process the 500 frames without any slowdown and very fast. Memory usage  stayed even below 10GB most of the time. Beta42 was taken 30-40 GB at some point..

I will build a test version for Windows and when it is ready, I will share it with you so you can test it quickly 😉

Mabula



   
ReplyQuote
(@mabula-admin)
Universe Admin
Joined: 9 years ago
Posts: 5369
 

Hi @wheeljack,

Please try this test version with a very different internal memory configuration. I did several tests and all were very solid with large datasets and fast, technically it definitely solves the problem if I analyzed it correctly. It could very well be that this new configuration is the way forward from here 😉

DOWNLOAD 2.0.0-beta43 ZGC test

This version should not stall/freeze at all like you experienced. And memory usage should not exceed 16GB of RAM by much unless processing GigaPixel images. This is what I see and how I configured it for this test release.

Look forward to hear about your testing with this one 😉

Mabula


This post was modified 4 months ago 2 times by Mabula-Admin

   
ReplyQuote
(@wheeljack)
Red Giant
Joined: 7 years ago
Posts: 40
Topic starter  

That was seriously quick response. Thanks alot. Glad to see that you also managed to spot the same (or at least similar) issues that I was. I will install your test-version and run the same dataset again, also with heavy LNC (to reproduce the same conditions) and report back to you. Most likely tomorrow. 

Thanks again for your efforts. It really is appreciated. 



   
ReplyQuote
(@wheeljack)
Red Giant
Joined: 7 years ago
Posts: 40
Topic starter  

Got the opportunity to test again, using all the same settings. 

Major improvements. Did not experience any stalls or freezes during either of the steps, until it came to LNC Calculations. The entire process (calibration, registration, normalization etc) took less than an hour to complete for all 724 frames, but it has been stuck here now for a little over 30 minutes: 

image

 

After reaching that stage, it went on for 5-10 minutes, then the process monitors died down, as well as my CPU fans. I let it just linger in the background while I was typing this, for another 20 minutes or so, and suddenly the CPU fans started up again and the process monitor indicated that something started to happen. That has now gone on for another 10 minutes and I am not sure if something is stuck somewhere or if it's actually doing anything. 

Prior to trying your new test version, I did try another run using Beta42 with 2nd degree LNC and all of those iterations finished after ~10m each. So might be something with 4th? 

I encountered no issues up until this point, so these are still major improvements. I let it work for another few minutes and cancelled after the first iteration of "performing LNC calculations" had ran for a total of 40 minutes. 

Attached the console-dump if it's helpful.

 

 



   
ReplyQuote
(@mabula-admin)
Universe Admin
Joined: 9 years ago
Posts: 5369
 

Hi @wheeljack,

Awesome ! I think this is all as expected now with regard to stability and APP being able to complete the long tasks. The new memory controls have fixed it. I have done more testing last night and all is very stable and fast now. The memory defragmentation is completely gone.

Calculating the LNC parameters 4th degree for more than 700 frames, does take a while. Part (33%) of the LNC calculation is still single threaded unfortunately so that explains the behaviour, that at some point the fans stop... and then once the single threaded part is done, the fans turn on again, and then the final part is calculated after which the result should be fine. LNC 2nd degree should be faster, less parameters to solve. And LNC 1st much faster.

Beta39 and earlier with the older LNC, was not even able to complete LNC1st degree I think on 700 images, or if it was, it took a huge amount of memory and time.

Thanks for testing ! I will release 2.0.0-beta43 with this fix for all platforms now. (and 2 other fixes).

Mabula



   
ReplyQuote
(@wheeljack)
Red Giant
Joined: 7 years ago
Posts: 40
Topic starter  

I can stick with 2nd degree LNC. I think that's what I have mostly been using for all these years anyway. I think I only started using 4th a few months ago out of curiosity, but on much smaller stacks (~200 ish) and have just been using it since. 

Thanks again for your work with this. Will giive Beta43 a run later today, and set it to 2nd degree LNC instead and continue to be using that (or experiment with differences between 2nd and 1st degree)



   
ReplyQuote
(@mabula-admin)
Universe Admin
Joined: 9 years ago
Posts: 5369
 

Hi @wheeljack,

Okay, great 😉 You are most welcome and thank you very much again for reporting this instability. How did beta43 run?

I will probably make another adjustment today or tomorrow with the new memory configuration, so beta44 can be released in the next day or 2 😉

Mabula



   
ReplyQuote
(@wheeljack)
Red Giant
Joined: 7 years ago
Posts: 40
Topic starter  

I got to try Beta43 a couple of times, and this is indeed a solid improvement. Even with 2nd degree LNC it did spend a lot of time on the "calculating" step, but as you already pointed out, that was most likely because of the size of the stack (700+ subs). I did do a few multi-NB stacks where I extracted HA and OIII (from a different set of subs, about 170 of them - same target) and did not encounter any issues at all with the stacking. The stacked image came out with horrible gradients (due to poor data) so I got to try the process several times in an effort to find and weed out the bad subs - and each run took only a few minutes to complete.



   
ReplyQuote
(@mabula-admin)
Universe Admin
Joined: 9 years ago
Posts: 5369
 

Hi @wheeljack,

Great that is exactly what we look for, stable, fast and continuous integration on the big tasks 😉 !

I have made further adjustments with the new memory controls in beta44, because it gave a couple of problems with systems with little memory and some warnings on the command line. I think that is all fixed with 2.0.0-beta44 now.

So please update to 2.0.0-beta44 and let me know how that goes.

Mabula



   
ReplyQuote
Share: