I am back? Sort of...
I decided today that I wish to re-establish my online presence. At least for now. Who knows how long this can last... This is a long story maybe. One that is sad to tell... ..
Once upon a time there was a homelab... One that consisted of not only a NAS but a PVE cluster that hosted MANY MANY webservices. Originally that cluster ran ZFS, and when I was initially testing the cluster features of PVE for the first time... Well I got the itch for some properly implemented HA via PVE. I found that ZFS was not a proper HA solution and so I assessed my options which were ceph zfs or external vsan or NAS etc and so based off my goals I immediately deployed CEPH. If you are familiar with the undelaying technology, you might be aware that the learning curve of CEPH can be a little bit more steep when you step beyond the most frequently traveled path...
Welp, put plainly. I wanted to use a feature in Windows 11 that allowed me to mount my RBD directly to my Windows PC so I could have Highly Available storage mounted as local storage on my main desktop and gaming PCs/Workstations. However I struggled to figure out how to adapt my networking virtually by use of a routing technology like FRR and I wished to avoid wireguard for this purpose. I wanted a truly local and near native solution. I consider FRR to be near native because practically everything under the sun supports it since the dawn of time networking wise at least (Windows doesn't ðŸ˜).
Somehow by a series of unfortunate events and I would say, user errors and also undoubtedly due to my inexperience with CEPH; I killed my cluster that was only comprised of 3 nodes (5-7 recommended, 3 the bare minimum) by changing the ip addresses to separate out the public and private bits of the cluster into 2 separate networks. The resulting network config against my better understanding was unrouteable and I did not go about changing the ips by the best method. I did not stop the cluster and ensure that no activity continued while this change was processed. In troubleshooting this I attempted to recover my monitor store when my PGs became unidentified and a host of other issues. I will be giving a guide on ceph that details everything I have learned here and to replace the snippet I published some time ago that laid out the logistics of how I decided the hardware for the cluster you are reading this from. I highly recommend to anyone reading this, tho I did post a guide on how to recover the store and likely I could have saved my data. But please! Just don't put yourself in that same position to start with! Trust me it's not something you want to do, I followed the official recovery methods. Theres quite a few resources linked in my github, CEPH recovery is a thing. ~ It may be possible to recover CEPH data when lost, but it is not a guarantee. Also the experts who can give it the best shot to be recovered don't come cheap.
The biggest mistake I made of all times here and one that I have no idea how I survived for 7 years without and this would be my first incident. I will be recommending this above all from now on... No joking around anymore. Backups are super important. You may have good practices but all it takes is for you to start getting yourself involved in newer technologies and implementing them into your production environment when you think you are ready to leave "testing" and you can follow in my footsteps. Please do not do that and create backups if you are reading this and share this same passion for homelab and homeserver, I hope no one who I come into contact with ever faces such a thing as I did. Data loss is no joke. I didn't lose my NAS. I have a lot that didn't get destroyed but I lost all my app data and every DB I had running.
The timing coincided with getting let go from another job and breaking it off in a relationship. I attribute the long time its taken my to get my webservices back online to overall depression. I lost motivation with everything hitting me all at once.
I think this was meant to happen. I have no regrets and think my lack of backups honestly justifies the loss. I deserved it. I am glad it happened to me. Especially now. I have aspirations to run a start up MSP. I can't afford to have issues like this when I am supporting clients so I am beyond glad that I was able to learn this lesson before I had any data that was more important without backup processes. Going forward I have a daily backup schedule for all of my app data and EVERY VM running in my cluster. And I am learning kubernetes and plan to learn flux so I can build out the next iteration of my homelab in git with the source replicated many places so if somehow by some odd chance I ever lose it again, even with backups lost if any replica of the code is retained I will be able to redeploy the environment fresh again and at least have a jump start into starting over. And also have the ability to easily document setup process of my lab for anyone curious.
I do not perceive that I was missed in my absence, I imagine no one missed my lab more than myself. However to anyone who did find the sites that I maintain down and was disappointed in their still continuing downtime. I would like to apologize and promise. It will be much better coming forward. And I will be keeping everyone posted with updates on how we will be breathing new life into our self hosted / home lab approaches!
If you read this, thank you for your time. I hope you found value in there somewhere. I can't imagine what part the value might be in. If you are curious about anything if you know how to reach me don't be afraid to reach out. Anyways. Everyone stay safe and keep your head up! Much love.
Best,
SoFMeRight