My HomeLab Monitoring Rules // feat. Checkmk
Read full transcript 14 segments
-
I love having full visibility into my I love having full visibility into my entire home lab infrastructure, entire home lab infrastructure, entire home lab infrastructure, monitoring servers, containers, Proxmox, monitoring servers, containers, Proxmox, monitoring servers, containers, Proxmox, storage, because I just want to know storage, because I just want to know storage, because I just want to know what's going on everywhere. But once you what's going on everywhere. But once you what's going on everywhere. But once you start to monitor everything, it also start to monitor everything, it also start to monitor everything, it also becomes very noisy. And at some point, becomes very noisy. And at some point, becomes very noisy. And at some point, you just have to decide what actually you just have to decide what actually you just have to decide what actually matters and what not. Otherwise, your matters and what not. Otherwise, your matters and what not. Otherwise, your dashboard slowly turns into a giant wall dashboard slowly turns into a giant wall dashboard slowly turns into a giant wall of yellow and red alerts. And after a of yellow and red alerts. And after a of yellow and red alerts. And after a while, you stop paying attention or you while, you stop paying attention or you while, you stop paying attention or you miss that one important alert. And miss that one important alert. And miss that one important alert. And honestly, I have to admit, this is honestly, I have to admit, this is honestly, I have to admit, this is exactly what happened in my own home exactly what happened in my own home exactly what happened in my own home lab, because I just hadn't tuned my lab, because I just hadn't tuned my lab, because I just hadn't tuned my monitoring setup for a while. But I've monitoring setup for a while. But I've monitoring setup for a while. But I've now solved that with Checkmk, the now solved that with Checkmk, the now solved that with Checkmk, the monitoring platform that gives you an monitoring platform that gives you an monitoring platform that gives you an incredible amount of visibility into incredible amount of visibility into incredible amount of visibility into your infrastructure right out of the your infrastructure right out of the your infrastructure right out of the box. And today, I'll show you how box. And today, I'll show you how box. And today, I'll show you how exactly I've tuned this setup to reduce exactly I've tuned this setup to reduce exactly I've tuned this setup to reduce the noise, adjusting thresholds, and the noise, adjusting thresholds, and the noise, adjusting thresholds, and building alerts that I can actually building alerts that I can actually building alerts that I can actually trust again. And so, in this video, trust again. And so, in this video, trust again. And so, in this video, let's dive deep into infrastructure let's dive deep into infrastructure let's dive deep into infrastructure monitoring. And a big thanks goes out to monitoring. And a big thanks goes out to monitoring. And a big thanks goes out to Checkmk for sponsoring this video.
-
Checkmk for sponsoring this video. Checkmk for sponsoring this video. All right, guys. So, note, this is not a All right, guys. So, note, this is not a All right, guys. So, note, this is not a full Checkmk installation or setup full Checkmk installation or setup full Checkmk installation or setup tutorial, because I've already made a tutorial, because I've already made a tutorial, because I've already made a full deep dive on Checkmk. Of course, full deep dive on Checkmk. Of course, full deep dive on Checkmk. Of course, I'll leave you that in the description. I'll leave you that in the description. I'll leave you that in the description. So, if you're completely new to Checkmk So, if you're completely new to Checkmk So, if you're completely new to Checkmk or infrastructure monitoring in general, or infrastructure monitoring in general, or infrastructure monitoring in general, then watch that one first. There's just then watch that one first. There's just then watch that one first. There's just one important update though, Checkmk one important update though, Checkmk one important update though, Checkmk released a new version 2.5, where the released a new version 2.5, where the released a new version 2.5, where the Docker image name changed from Checkmk Docker image name changed from Checkmk Docker image name changed from Checkmk raw to Checkmk community. So, be aware raw to Checkmk community. So, be aware raw to Checkmk community. So, be aware of that. Of course, I've already made an of that. Of course, I've already made an of that. Of course, I've already made an update to my boilerplate repository, so update to my boilerplate repository, so update to my boilerplate repository, so that is what you can use to easily spin that is what you can use to easily spin that is what you can use to easily spin up a new Checkmk installation using up a new Checkmk installation using up a new Checkmk installation using Docker and Docker Compose. And because Docker and Docker Compose. And because Docker and Docker Compose. And because I'm using Checkmk 2.5 in this video, I'm using Checkmk 2.5 in this video, I'm using Checkmk 2.5 in this video, some of the menus may look a bit some of the menus may look a bit some of the menus may look a bit different than my older tutorials and different than my older tutorials and different than my older tutorials and throughout the video we'll also use some throughout the video we'll also use some throughout the video we'll also use some of the new features like the global of the new features like the global of the new features like the global search function or background activation search function or background activation search function or background activation which makes tuning rules and working which makes tuning rules and working which makes tuning rules and working with Checkmk a lot easier and faster. with Checkmk a lot easier and faster. with Checkmk a lot easier and faster. Now, before we go over some specific Now, before we go over some specific Now, before we go over some specific examples of the Checkmk monitoring rules examples of the Checkmk monitoring rules examples of the Checkmk monitoring rules and what I've changed in my setup, let and what I've changed in my setup, let and what I've changed in my setup, let us quickly recap how Checkmk works and us quickly recap how Checkmk works and us quickly recap how Checkmk works and let me explain how I'm managing my let me explain how I'm managing my let me explain how I'm managing my monitoring environment. Because Checkmk monitoring environment. Because Checkmk monitoring environment. Because Checkmk does not just monitor any hosts, it also does not just monitor any hosts, it also does not just monitor any hosts, it also discovers services on them and then you discovers services on them and then you discovers services on them and then you have all of these rules that define how have all of these rules that define how have all of these rules that define how these services behave. All of these these services behave. All of these these services behave. All of these rules can match to a host name, a host rules can match to a host name, a host rules can match to a host name, a host tag, a host label, or service names and
-
tag, a host label, or service names and tag, a host label, or service names and service labels. And this is where I service labels. And this is where I service labels. And this is where I believe a lot of people make monitoring believe a lot of people make monitoring believe a lot of people make monitoring harder than it actually needs to be harder than it actually needs to be harder than it actually needs to be because if you set up all of these rules because if you set up all of these rules because if you set up all of these rules and attach them to all of the individual and attach them to all of the individual and attach them to all of the individual hosts or services on your system, then hosts or services on your system, then hosts or services on your system, then your monitoring setup becomes a pile of your monitoring setup becomes a pile of your monitoring setup becomes a pile of one-off exceptions that doesn't really one-off exceptions that doesn't really one-off exceptions that doesn't really scale well even in a home lab. And a scale well even in a home lab. And a scale well even in a home lab. And a really important tip that I'd like to really important tip that I'd like to really important tip that I'd like to share with you is when you start setting share with you is when you start setting share with you is when you start setting up monitoring rules, first of all, up monitoring rules, first of all, up monitoring rules, first of all, attach labels to all of your Checkmk attach labels to all of your Checkmk attach labels to all of your Checkmk hosts. For example, here I'm managing hosts. For example, here I'm managing hosts. For example, here I'm managing all of my hosts like the server port 1 2 all of my hosts like the server port 1 2 all of my hosts like the server port 1 2 3, my virtual machines as well as 3, my virtual machines as well as 3, my virtual machines as well as Proxmox servers. And as you can see Proxmox servers. And as you can see Proxmox servers. And as you can see here, I'm attaching two labels, the type here, I'm attaching two labels, the type here, I'm attaching two labels, the type and the platform. For example, here it's and the platform. For example, here it's and the platform. For example, here it's Proxmox, the platform is bare metal. Proxmox, the platform is bare metal. Proxmox, the platform is bare metal. Here on this server, it's type server, Here on this server, it's type server, Here on this server, it's type server, the platform is VM. And I'm also the platform is VM. And I'm also the platform is VM. And I'm also managing several Docker containers as managing several Docker containers as managing several Docker containers as hosts to monitor all of the service hosts to monitor all of the service hosts to monitor all of the service applications as well. And these usually applications as well. And these usually applications as well. And these usually get the label type compose and the get the label type compose and the get the label type compose and the platform Docker because I can later say platform Docker because I can later say platform Docker because I can later say in all of the rule settings, "Hey, in all of the rule settings, "Hey, in all of the rule settings, "Hey, please attach this rule only to hosts please attach this rule only to hosts please attach this rule only to hosts that have the label type server or type that have the label type server or type that have the label type server or type Proxmox and not just say, "Hey, create a Proxmox and not just say, "Hey, create a Proxmox and not just say, "Hey, create a new rule for server prod one." and then new rule for server prod one." and then new rule for server prod one." and then copy that rule to server prod two and copy that rule to server prod two and copy that rule to server prod two and then later on you forget about server then later on you forget about server then later on you forget about server prod three and so on. So, this is really prod three and so on. So, this is really prod three and so on. So, this is really important. Turned my Checkmk rules from important. Turned my Checkmk rules from important. Turned my Checkmk rules from fragile host specific exceptions into
-
fragile host specific exceptions into fragile host specific exceptions into reusable monitoring policies that I can reusable monitoring policies that I can reusable monitoring policies that I can really attach to specific groups of really attach to specific groups of really attach to specific groups of hosts. What I also love to do is I love hosts. What I also love to do is I love hosts. What I also love to do is I love to use the Checkmk UI for exploring the to use the Checkmk UI for exploring the to use the Checkmk UI for exploring the different rules because this is honestly different rules because this is honestly different rules because this is honestly much faster when you're just trying to much faster when you're just trying to much faster when you're just trying to figure out things. But once a rule is figure out things. But once a rule is figure out things. But once a rule is correct, I want automation. So, that's correct, I want automation. So, that's correct, I want automation. So, that's why my real Checkmk rules also live in why my real Checkmk rules also live in why my real Checkmk rules also live in Ansible YAML files here. And that gives Ansible YAML files here. And that gives Ansible YAML files here. And that gives me two things. I can reproduce this me two things. I can reproduce this me two things. I can reproduce this setup easily and I can also remember why setup easily and I can also remember why setup easily and I can also remember why a rule exists six months later. And a rule exists six months later. And a rule exists six months later. And because these rules are basically just because these rules are basically just because these rules are basically just simple YAML files, I can even let an AI simple YAML files, I can even let an AI simple YAML files, I can even let an AI agent help me to review all of these agent help me to review all of these agent help me to review all of these rules and adjust or change certain rules and adjust or change certain rules and adjust or change certain settings instead of just me having to go settings instead of just me having to go settings instead of just me having to go through the web UI and changing things through the web UI and changing things through the web UI and changing things by hand. This is by the way also where by hand. This is by the way also where by hand. This is by the way also where Checkmk 2.5 fits nicely into this Checkmk 2.5 fits nicely into this Checkmk 2.5 fits nicely into this workflow because if you're exploring workflow because if you're exploring workflow because if you're exploring things, you can easily use the global things, you can easily use the global things, you can easily use the global search function. For example, if you're search function. For example, if you're search function. For example, if you're searching for a rule that sets up a searching for a rule that sets up a searching for a rule that sets up a temperature setting on your temperature setting on your temperature setting on your infrastructure without having to infrastructure without having to infrastructure without having to remember where everything lives in the remember where everything lives in the remember where everything lives in the menu. That's really practical. And once menu. That's really practical. And once menu. That's really practical. And once I'm happy, I put the rule in Ansible.
-
I'm happy, I put the rule in Ansible. I'm happy, I put the rule in Ansible. So, this is my preferred workflow. By So, this is my preferred workflow. By So, this is my preferred workflow. By the way, if you want to learn more about the way, if you want to learn more about the way, if you want to learn more about how to use Ansible to automate your how to use Ansible to automate your how to use Ansible to automate your Checkmk setup, I've also made a full Checkmk setup, I've also made a full Checkmk setup, I've also made a full tutorial on that. So, check this out if tutorial on that. So, check this out if tutorial on that. So, check this out if you like. I'll put you this video in the you like. I'll put you this video in the you like. I'll put you this video in the description as well. But just to make it description as well. But just to make it description as well. But just to make it clear, you don't necessarily have to use clear, you don't necessarily have to use clear, you don't necessarily have to use Ansible to configure Checkmk. You, for Ansible to configure Checkmk. You, for Ansible to configure Checkmk. You, for example, could also use web UI. This is example, could also use web UI. This is example, could also use web UI. This is totally fine as well. So, whatever fits totally fine as well. So, whatever fits totally fine as well. So, whatever fits your workflow best, I'm not saying that your workflow best, I'm not saying that your workflow best, I'm not saying that you have to follow my best practices. you have to follow my best practices. you have to follow my best practices. But anyways, whether you create the But anyways, whether you create the But anyways, whether you create the rules in the web interface or if you're rules in the web interface or if you're rules in the web interface or if you're using Ansible, you're basically always using Ansible, you're basically always using Ansible, you're basically always facing the same tasks. Normally, you facing the same tasks. Normally, you facing the same tasks. Normally, you just let checkmk discover any services just let checkmk discover any services just let checkmk discover any services on any of your hosts in your on any of your hosts in your on any of your hosts in your infrastructure and you just accept all infrastructure and you just accept all infrastructure and you just accept all of this and then it automatically starts of this and then it automatically starts of this and then it automatically starts monitoring all of these different monitoring all of these different monitoring all of these different services on the system. But, and this is services on the system. But, and this is services on the system. But, and this is important, not every discovered service important, not every discovered service important, not every discovered service deserves the same level of attention. deserves the same level of attention. deserves the same level of attention. The more hosts you will add to your The more hosts you will add to your The more hosts you will add to your monitoring setup, the more service monitoring setup, the more service monitoring setup, the more service notifications will pop up in the notifications will pop up in the notifications will pop up in the dashboard reporting problems or things dashboard reporting problems or things dashboard reporting problems or things you should take care of. And then you you should take care of. And then you you should take care of. And then you need to decide, is this actually an need to decide, is this actually an need to decide, is this actually an important alert or is this just noise?
-
important alert or is this just noise? important alert or is this just noise? And this is where monitoring gets very And this is where monitoring gets very And this is where monitoring gets very personal. Let me make this clear. What personal. Let me make this clear. What personal. Let me make this clear. What matters in my home lab doesn't matters in my home lab doesn't matters in my home lab doesn't automatically have to matter in yours. automatically have to matter in yours. automatically have to matter in yours. So, treat all of the following examples So, treat all of the following examples So, treat all of the following examples just as an inspiration, not just as just as an inspiration, not just as just as an inspiration, not just as copy-paste rules. Because the wrong fix copy-paste rules. Because the wrong fix copy-paste rules. Because the wrong fix is just mute or disable all of the is just mute or disable all of the is just mute or disable all of the critical services. The better approach critical services. The better approach critical services. The better approach is to actually tune your monitoring is to actually tune your monitoring is to actually tune your monitoring system so you only get those alerts that system so you only get those alerts that system so you only get those alerts that matter in your setup. The goal is not to matter in your setup. The goal is not to matter in your setup. The goal is not to make every dashboard green. The goal is make every dashboard green. The goal is make every dashboard green. The goal is to make the dashboard useful. So, let me to make the dashboard useful. So, let me to make the dashboard useful. So, let me show you one example. Let us go into show you one example. Let us go into show you one example. Let us go into setup, services, and then go to setup, services, and then go to setup, services, and then go to discovery rules. So, here you can see discovery rules. So, here you can see discovery rules. So, here you can see one rule, disabled services. So, here one rule, disabled services. So, here one rule, disabled services. So, here I've added three separate services that I've added three separate services that I've added three separate services that are attached to specific host labels. are attached to specific host labels. are attached to specific host labels. Remember, these are the groups of all of Remember, these are the groups of all of Remember, these are the groups of all of the different hosts like virtual the different hosts like virtual the different hosts like virtual machines or Proxmox infrastructure. machines or Proxmox infrastructure. machines or Proxmox infrastructure. First of all, let me show you this rule First of all, let me show you this rule First of all, let me show you this rule here, the disabled K3s dynamic file here, the disabled K3s dynamic file here, the disabled K3s dynamic file system monitoring. Because this was one system monitoring. Because this was one system monitoring. Because this was one one the most annoying examples that one the most annoying examples that one the most annoying examples that always popped up in my notification. The always popped up in my notification. The always popped up in my notification. The problem is that Checkmk discovers any problem is that Checkmk discovers any problem is that Checkmk discovers any file system mount point as separate file system mount point as separate file system mount point as separate services. So, all of these different services. So, all of these different services. So, all of these different paths show up as different service paths show up as different service paths show up as different service entries, but treating them like a normal entries, but treating them like a normal entries, but treating them like a normal file system is not really useful for me file system is not really useful for me file system is not really useful for me because these are runtime mount points because these are runtime mount points because these are runtime mount points for containers. And it is really for containers. And it is really for containers. And it is really expected that in a Kubernetes cluster, expected that in a Kubernetes cluster, expected that in a Kubernetes cluster, many of these containers appear, many of these containers appear, many of these containers appear, disappear, or change as workloads move
-
disappear, or change as workloads move disappear, or change as workloads move around. And when you don't disable these around. And when you don't disable these around. And when you don't disable these type of services, you always see Checkmk type of services, you always see Checkmk type of services, you always see Checkmk popping up these alerts. Not exactly popping up these alerts. Not exactly popping up these alerts. Not exactly that kind of alert that you want to see that kind of alert that you want to see that kind of alert that you want to see in your dashboard. So, therefore I added in your dashboard. So, therefore I added in your dashboard. So, therefore I added this rule to ignore this. And here you this rule to ignore this. And here you this rule to ignore this. And here you can see that's a condition match on can see that's a condition match on can see that's a condition match on anything that starts with file anything that starts with file anything that starts with file system/run/k3s/containerd system/run/k3s/containerd system/run/k3s/containerd and anything that comes after that. So, and anything that comes after that. So, and anything that comes after that. So, that matches all of these random that matches all of these random that matches all of these random container file system paths. The next container file system paths. The next container file system paths. The next example of disabled services is this example of disabled services is this example of disabled services is this here. The Proxmox VE memory usage. Here here. The Proxmox VE memory usage. Here here. The Proxmox VE memory usage. Here on this virtual machine, for example, on this virtual machine, for example, on this virtual machine, for example, this would have produced a warning this would have produced a warning this would have produced a warning message in the dashboard because the message in the dashboard because the message in the dashboard because the usage of the virtual memory is 88%. usage of the virtual memory is 88%. usage of the virtual memory is 88%. That's above the threshold of warnings That's above the threshold of warnings That's above the threshold of warnings that starts at 80%. When it's above 90%, that starts at 80%. When it's above 90%, that starts at 80%. When it's above 90%, you even get a critical warning in your you even get a critical warning in your you even get a critical warning in your monitoring setup. This, by the way, you monitoring setup. This, by the way, you monitoring setup. This, by the way, you could also see on your Proxmox system. could also see on your Proxmox system. could also see on your Proxmox system. Here the memory usage is also 88-89%. Here the memory usage is also 88-89%. Here the memory usage is also 88-89%. the same that Checkmk discovers. But the same that Checkmk discovers. But the same that Checkmk discovers. But when we log into the actual virtual when we log into the actual virtual when we log into the actual virtual machine and take a look at what is the machine and take a look at what is the machine and take a look at what is the real memory usage, we get a completely real memory usage, we get a completely real memory usage, we get a completely different value. So, this even looks different value. So, this even looks different value. So, this even looks from a Proxmox side of view much more from a Proxmox side of view much more from a Proxmox side of view much more dramatic than it actually is on the dramatic than it actually is on the dramatic than it actually is on the operating system itself. And that is operating system itself. And that is operating system itself. And that is because from the Proxmox side, a Linux because from the Proxmox side, a Linux because from the Proxmox side, a Linux VM can often look like it's using almost VM can often look like it's using almost VM can often look like it's using almost all of its memory because Linux uses all of its memory because Linux uses all of its memory because Linux uses available RAM for cache and buffers, and available RAM for cache and buffers, and available RAM for cache and buffers, and the hypervisor only sees the VM mostly
-
the hypervisor only sees the VM mostly the hypervisor only sees the VM mostly from the outside, but it does not really from the outside, but it does not really from the outside, but it does not really automatically mean the guest is under automatically mean the guest is under automatically mean the guest is under memory pressure. CheckMK gives us the memory pressure. CheckMK gives us the memory pressure. CheckMK gives us the much better metric. You can also see much better metric. You can also see much better metric. You can also see this here under memory. Total virtual this here under memory. Total virtual this here under memory. Total virtual memory is 62%, so nearly the same that memory is 62%, so nearly the same that memory is 62%, so nearly the same that we just discovered on the operating we just discovered on the operating we just discovered on the operating system. This is why in my monitoring system. This is why in my monitoring system. This is why in my monitoring setup, I just ignore this service that setup, I just ignore this service that setup, I just ignore this service that just sees the memory from the hypervisor just sees the memory from the hypervisor just sees the memory from the hypervisor side of view, and I use the actual side of view, and I use the actual side of view, and I use the actual memory service to better see what a VM memory service to better see what a VM memory service to better see what a VM is actually doing or what the resources is actually doing or what the resources is actually doing or what the resources are on that virtual machine. The same are on that virtual machine. The same are on that virtual machine. The same idea I'm also following for the Proxmox idea I'm also following for the Proxmox idea I'm also following for the Proxmox LVM volume groups, because a Proxmox LVM volume groups, because a Proxmox LVM volume groups, because a Proxmox volume group can also look almost full volume group can also look almost full volume group can also look almost full allocated. Note, this is not for the allocated. Note, this is not for the allocated. Note, this is not for the virtual machines, this is for the virtual machines, this is for the virtual machines, this is for the Proxmox hypervisor itself. So, if we Proxmox hypervisor itself. So, if we Proxmox hypervisor itself. So, if we just take a look at the disks LVM, you just take a look at the disks LVM, you just take a look at the disks LVM, you can see this is completely full to 100%. can see this is completely full to 100%. can see this is completely full to 100%. The second hard drive is 98%, so this The second hard drive is 98%, so this The second hard drive is 98%, so this also looks very dramatic, but of course, also looks very dramatic, but of course, also looks very dramatic, but of course, that does not automatically mean that that does not automatically mean that that does not automatically mean that the storage I care about is actually the storage I care about is actually the storage I care about is actually full or that the VM is about to run out full or that the VM is about to run out full or that the VM is about to run out of disk space. So, for troubleshooting of disk space. So, for troubleshooting of disk space. So, for troubleshooting or alerting, I use another metric that or alerting, I use another metric that or alerting, I use another metric that is much better for this. If we go to the is much better for this. If we go to the is much better for this. If we go to the Proxmox machine itself here on the LVM Proxmox machine itself here on the LVM Proxmox machine itself here on the LVM group PVE data, the data is using 80%, group PVE data, the data is using 80%, group PVE data, the data is using 80%, so this is far away the 98 or 100% so this is far away the 98 or 100% so this is far away the 98 or 100% usage. So, just add a new rule that
-
usage. So, just add a new rule that usage. So, just add a new rule that matches all your Proxmox machines, matches all your Proxmox machines, matches all your Proxmox machines, disabling services that start with disabling services that start with disabling services that start with LVMVG, and then the {dot} asterisk, so LVMVG, and then the {dot} asterisk, so LVMVG, and then the {dot} asterisk, so anything that comes after that. This anything that comes after that. This anything that comes after that. This helps to automatically ignore not just helps to automatically ignore not just helps to automatically ignore not just the LVM local, also any Ceph storage the LVM local, also any Ceph storage the LVM local, also any Ceph storage that starts with LVM volume group. And that starts with LVM volume group. And that starts with LVM volume group. And another service that I also to another service that I also to another service that I also to let me just go back to the setup let me just go back to the setup let me just go back to the setup discovery rule. And this is in the discovery rule. And this is in the discovery rule. And this is in the network interface and switch port network interface and switch port network interface and switch port discovery. So, here I've added a rule on discovery. So, here I've added a rule on discovery. So, here I've added a rule on all of my machines to just ignore any all of my machines to just ignore any all of my machines to just ignore any networks that starts with docker zero or networks that starts with docker zero or networks that starts with docker zero or docker underscore anything bridge docker underscore anything bridge docker underscore anything bridge interfaces, tap interfaces. Of course, interfaces, tap interfaces. Of course, interfaces, tap interfaces. Of course, Checkmk will notify you when an Checkmk will notify you when an Checkmk will notify you when an interface goes down, so you get an interface goes down, so you get an interface goes down, so you get an alert. And for these temporary alert. And for these temporary alert. And for these temporary interfaces, I don't want to be interfaces, I don't want to be interfaces, I don't want to be distracted with all of the alerts of distracted with all of the alerts of distracted with all of the alerts of services that I actually don't care services that I actually don't care services that I actually don't care about. Of course, these are just a very about. Of course, these are just a very about. Of course, these are just a very few examples. I've created some other few examples. I've created some other few examples. I've created some other rules like this, but I think I don't rules like this, but I think I don't rules like this, but I think I don't have to go and show you every single have to go and show you every single have to go and show you every single rule. I think monitoring is very rule. I think monitoring is very rule. I think monitoring is very personally. So, don't just copy my personally. So, don't just copy my personally. So, don't just copy my rules. Better understand how to actually rules. Better understand how to actually rules. Better understand how to actually write them yourself that fit your write them yourself that fit your write them yourself that fit your specific environment, not just mine. But specific environment, not just mine. But specific environment, not just mine. But speaking of that, let us also go through speaking of that, let us also go through speaking of that, let us also go through the second group of alerts that is a bit the second group of alerts that is a bit the second group of alerts that is a bit different than just disabling stuff. And different than just disabling stuff. And different than just disabling stuff. And honestly, it might also be a bit more honestly, it might also be a bit more honestly, it might also be a bit more difficult to researching. But these are difficult to researching. But these are difficult to researching. But these are checks that I actually care about. I checks that I actually care about. I checks that I actually care about. I just don't care about them at the just don't care about them at the just don't care about them at the default thresholds. Because instead of default thresholds. Because instead of default thresholds. Because instead of just disabling a good check, we can use
-
just disabling a good check, we can use just disabling a good check, we can use Checkmk rules to teach the monitoring Checkmk rules to teach the monitoring Checkmk rules to teach the monitoring setup of what normal means for this setup of what normal means for this setup of what normal means for this specific host or service. So, one really specific host or service. So, one really specific host or service. So, one really good example of alerts that always good example of alerts that always good example of alerts that always annoyed me is the committed memory on annoyed me is the committed memory on annoyed me is the committed memory on virtual servers. That virtual machine virtual servers. That virtual machine virtual servers. That virtual machine has a total virtual memory of about 23 has a total virtual memory of about 23 has a total virtual memory of about 23 GB bytes, but the committed memory is GB bytes, but the committed memory is GB bytes, but the committed memory is much higher than the total virtual much higher than the total virtual much higher than the total virtual memory. And this is because the memory. And this is because the memory. And this is because the committed memory is not the same thing committed memory is not the same thing committed memory is not the same thing as the RAM is full. On Linux, as the RAM is full. On Linux, as the RAM is full. On Linux, applications can reserve or promise applications can reserve or promise applications can reserve or promise memory that they are not actively using memory that they are not actively using memory that they are not actively using right now. So, especially on virtual right now. So, especially on virtual right now. So, especially on virtual machines, the committed memory can go machines, the committed memory can go machines, the committed memory can go way above 100% without the system way above 100% without the system way above 100% without the system immediately being under real memory immediately being under real memory immediately being under real memory pressure. And this is where the default pressure. And this is where the default pressure. And this is where the default values of checkmk are, in my opinion, a values of checkmk are, in my opinion, a values of checkmk are, in my opinion, a bit too sensitive. So, I think they bit too sensitive. So, I think they bit too sensitive. So, I think they start at 100% for warning, and then 150% start at 100% for warning, and then 150% start at 100% for warning, and then 150% for critical warnings. And as you've for critical warnings. And as you've for critical warnings. And as you've seen in these examples on almost all my seen in these examples on almost all my seen in these examples on almost all my virtual servers, I was getting yellow virtual servers, I was getting yellow virtual servers, I was getting yellow warnings. But, of course, I cannot just warnings. But, of course, I cannot just warnings. But, of course, I cannot just disable the memory service, yeah?
-
disable the memory service, yeah? disable the memory service, yeah? Because then I don't get memory warnings Because then I don't get memory warnings Because then I don't get memory warnings at all. And I don't want to ignore this. at all. And I don't want to ignore this. at all. And I don't want to ignore this. So, I still care about if the VM really So, I still care about if the VM really So, I still care about if the VM really starts running out of memory. I created starts running out of memory. I created starts running out of memory. I created a rule that sets the upper levels for a rule that sets the upper levels for a rule that sets the upper levels for committed memory are increased to let's committed memory are increased to let's committed memory are increased to let's start warning at 120%, and let's leave start warning at 120%, and let's leave start warning at 120%, and let's leave critical warnings at 150% of RAM plus critical warnings at 150% of RAM plus critical warnings at 150% of RAM plus swap. And I'm attaching this to all of swap. And I'm attaching this to all of swap. And I'm attaching this to all of the hosts that have the virtual machine the hosts that have the virtual machine the hosts that have the virtual machine platform label. So, this keeps all of my platform label. So, this keeps all of my platform label. So, this keeps all of my memory checks useful because it stops memory checks useful because it stops memory checks useful because it stops warning too early when it's above 100%, warning too early when it's above 100%, warning too early when it's above 100%, but it's still alerts me when the but it's still alerts me when the but it's still alerts me when the committed memory becomes too high above committed memory becomes too high above committed memory becomes too high above 120% or 150% so that I can actually 120% or 150% so that I can actually 120% or 150% so that I can actually investigate problems with the committed investigate problems with the committed investigate problems with the committed memory on that machine. Another good memory on that machine. Another good memory on that machine. Another good example is the NVMe temperature on all example is the NVMe temperature on all example is the NVMe temperature on all of my bare metal or physical machines. of my bare metal or physical machines. of my bare metal or physical machines. So, here for example, you can see I have So, here for example, you can see I have So, here for example, you can see I have services that monitor the temperature of services that monitor the temperature of services that monitor the temperature of the internal hard drives. So, I think by the internal hard drives. So, I think by the internal hard drives. So, I think by default checkmk produces a warning if default checkmk produces a warning if default checkmk produces a warning if the temperature goes for internal drive the temperature goes for internal drive the temperature goes for internal drive higher than 35° higher than 35° higher than 35° and a critical warning if it goes above and a critical warning if it goes above and a critical warning if it goes above 40°. And if you see the graphs here, you 40°. And if you see the graphs here, you 40°. And if you see the graphs here, you can clearly see that my internal hard can clearly see that my internal hard can clearly see that my internal hard drives are always above 40°, so this drives are always above 40°, so this drives are always above 40°, so this would always produce a critical warning would always produce a critical warning would always produce a critical warning on all of my physical servers. But, it on all of my physical servers. But, it on all of my physical servers. But, it is still not a problem in my setup is still not a problem in my setup is still not a problem in my setup because I'm using NVMe's, and NVMe
-
because I'm using NVMe's, and NVMe because I'm using NVMe's, and NVMe drives can get much warmer than normal drives can get much warmer than normal drives can get much warmer than normal hard drives, especially in a small hard drives, especially in a small hard drives, especially in a small server behind a GPU or under a heat sink server behind a GPU or under a heat sink server behind a GPU or under a heat sink with limited airflow. So, a temperature with limited airflow. So, a temperature with limited airflow. So, a temperature that looks scary to a magnetic hard that looks scary to a magnetic hard that looks scary to a magnetic hard drive or maybe even an SSD is completely drive or maybe even an SSD is completely drive or maybe even an SSD is completely normal for an NVMe drive. Still, because normal for an NVMe drive. Still, because normal for an NVMe drive. Still, because Checkmk doesn't differentiate between Checkmk doesn't differentiate between Checkmk doesn't differentiate between the different drive types, it will the different drive types, it will the different drive types, it will always produce a critical warning. So, always produce a critical warning. So, always produce a critical warning. So, therefore, I created a rule to adjust therefore, I created a rule to adjust therefore, I created a rule to adjust the upper temperature limits start a the upper temperature limits start a the upper temperature limits start a warning at 60°, critical at 85°. This is warning at 60°, critical at 85°. This is warning at 60°, critical at 85°. This is where even an NVMe starts getting hot where even an NVMe starts getting hot where even an NVMe starts getting hot and it needs to throttle down or it and it needs to throttle down or it and it needs to throttle down or it becomes unstable. So, I need to becomes unstable. So, I need to becomes unstable. So, I need to investigate this. And here, under the investigate this. And here, under the investigate this. And here, under the sensor ID, I added the conditions that sensor ID, I added the conditions that sensor ID, I added the conditions that if a sensor starts with /dev/nvme, if a sensor starts with /dev/nvme, if a sensor starts with /dev/nvme, the asterisk for anything that comes the asterisk for anything that comes the asterisk for anything that comes after that, or for specific names of the after that, or for specific names of the after that, or for specific names of the NVMe drive models. And then you can also NVMe drive models. And then you can also NVMe drive models. And then you can also see that Checkmk automatically adjusts see that Checkmk automatically adjusts see that Checkmk automatically adjusts these temperature warning and critical these temperature warning and critical these temperature warning and critical warnings in the graph view and it warnings in the graph view and it warnings in the graph view and it doesn't produce any false positives for doesn't produce any false positives for doesn't produce any false positives for NVMe critical temperature. These are NVMe critical temperature. These are NVMe critical temperature. These are some of the examples that I thought some of the examples that I thought some of the examples that I thought might be worth showing. Again, what might be worth showing. Again, what might be worth showing. Again, what matters in my home lab infrastructure matters in my home lab infrastructure matters in my home lab infrastructure doesn't necessarily have to matter in doesn't necessarily have to matter in doesn't necessarily have to matter in your infrastructure. I can just give you your infrastructure. I can just give you your infrastructure. I can just give you the tip, go to your monitoring hosts, the tip, go to your monitoring hosts, the tip, go to your monitoring hosts, see what services are producing yellow see what services are producing yellow see what services are producing yellow or red warnings, and then you just have or red warnings, and then you just have or red warnings, and then you just have to go through all of them and find out to go through all of them and find out to go through all of them and find out if they are reasonable warnings or
-
if they are reasonable warnings or if they are reasonable warnings or reasonable alerts or if you need to reasonable alerts or if you need to reasonable alerts or if you need to disable certain services or adjust disable certain services or adjust disable certain services or adjust thresholds in your setup and create thresholds in your setup and create thresholds in your setup and create label groups for the same type of hosts label groups for the same type of hosts label groups for the same type of hosts in your setup and attach these rules not in your setup and attach these rules not in your setup and attach these rules not only for one specific host, but attach only for one specific host, but attach only for one specific host, but attach them to the same type of host groups. them to the same type of host groups. them to the same type of host groups. And the one remaining puzzle piece here And the one remaining puzzle piece here And the one remaining puzzle piece here is to keeping the setup clean over time is to keeping the setup clean over time is to keeping the setup clean over time because home labs are not static. In a because home labs are not static. In a because home labs are not static. In a home lab, there are many, many things home lab, there are many, many things home lab, there are many, many things that change frequently like virtual that change frequently like virtual that change frequently like virtual machines that appear or disappear, uh machines that appear or disappear, uh machines that appear or disappear, uh services that you might deploy and then services that you might deploy and then services that you might deploy and then destroy the next week or so. I can just destroy the next week or so. I can just destroy the next week or so. I can just speak for myself. I regularly rebuild speak for myself. I regularly rebuild speak for myself. I regularly rebuild VMs, adding or removing containers, VMs, adding or removing containers, VMs, adding or removing containers, services, renaming things. So, that's services, renaming things. So, that's services, renaming things. So, that's why I usually also configure a periodic why I usually also configure a periodic why I usually also configure a periodic service discovery. So, just go into service discovery. So, just go into service discovery. So, just go into setup services discovery rules and then setup services discovery rules and then setup services discovery rules and then look for periodic service discovery. look for periodic service discovery. look for periodic service discovery. Here I configure things like Here I configure things like Here I configure things like automatically update service automatically update service automatically update service configuration, monitor undecided configuration, monitor undecided configuration, monitor undecided services, remove vanished services, and services, remove vanished services, and services, remove vanished services, and update host labels. You could also update host labels. You could also update host labels. You could also update service labels or service update service labels or service update service labels or service parameters, but I leave them off because parameters, but I leave them off because parameters, but I leave them off because I want a little more control in that I want a little more control in that I want a little more control in that area. And then perform a service area. And then perform a service area. And then perform a service discovery every hour. So, that means discovery every hour. So, that means discovery every hour. So, that means whenever I'm removing a service on a whenever I'm removing a service on a whenever I'm removing a service on a host or I'm changing an interface or host or I'm changing an interface or host or I'm changing an interface or whatever, this automatically monitors whatever, this automatically monitors whatever, this automatically monitors any undecided services that appear in any undecided services that appear in any undecided services that appear in the service discovery and remove any of the service discovery and remove any of the service discovery and remove any of the vanished services. I think this is the vanished services. I think this is the vanished services. I think this is really perfect for a home lab where
-
really perfect for a home lab where really perfect for a home lab where services change all the time. So, yeah, services change all the time. So, yeah, services change all the time. So, yeah, that's basically how I manage my Checkmk that's basically how I manage my Checkmk that's basically how I manage my Checkmk monitoring rules in my home lab. I hope monitoring rules in my home lab. I hope monitoring rules in my home lab. I hope this video helped you to get your this video helped you to get your this video helped you to get your monitoring system under control or at monitoring system under control or at monitoring system under control or at least get some inspiration of how to set least get some inspiration of how to set least get some inspiration of how to set up your monitoring rules, adjusting up your monitoring rules, adjusting up your monitoring rules, adjusting certain thresholds, and so on. So, thank certain thresholds, and so on. So, thank certain thresholds, and so on. So, thank you so much for watching. A big thanks you so much for watching. A big thanks you so much for watching. A big thanks goes out to all of my supporters and goes out to all of my supporters and goes out to all of my supporters and Checkmk for sponsoring this video. And Checkmk for sponsoring this video. And Checkmk for sponsoring this video. And of course, I'm going to catch you in the of course, I'm going to catch you in the of course, I'm going to catch you in the next one. Take care. Bye-bye.
Summary
The main theme is managing the noise and alert fatigue in home lab infrastructure monitoring. Key subjects include Proxmox, servers, containers, storage, and the Checkmk monitoring platform, specifically its version 2.5 and Docker Compose usage. The practical takeaway is that by tuning monitoring thresholds and building reliable alerts, you can regain trust in your dashboard and avoid missing critical issues, as demonstrated with Checkmk.