I am not the biggest fan of docker, using docker is like removing a molehill with a hand grenade. It's existance is indicative of the quality of the worlds software ecosystem as a whole, that people are willing to suffer large memory overheads and awful configuration processes; All because software is so awful that fixing it is no longer considered a viable option. Docker represents the rot that has overtaken the industry in the last decade.
That being said it is quite useful for deploying databases, so lets deploy Oracle on it. Oracle has a really good single instance docker image database/free which took me less than 10 minutes to get running with when I wanted to test my OCI (Oracle Call Interface not Oracle Cloud Infrastructure) db abstraction. Surely the RAC equivalent wouldn't be much harder to deploy, they even have a prebuilt image, a guide, and its open source on GitHub. I should have this done in an afternoon.
I opened the RAC docker compose guide and felt my heart sink a little.
#---------Bring up DNS------------
docker compose up -d ${DNS_CONTAINER_NAME} && docker compose stop ${DNS_CONTAINER_NAME}
docker network disconnect ${PUBLIC_NETWORK_NAME} ${DNS_CONTAINER_NAME}
docker network disconnect ${PRIVATE_NETWORK_NAME} ${DNS_CONTAINER_NAME}
docker network connect ${PUBLIC_NETWORK_NAME} --ip ${DNS_PUBLIC_IP} ${DNS_CONTAINER_NAME}
docker network connect ${PRIVATE_NETWORK_NAME} --ip ${DNS_PRIVATE_IP} ${DNS_CONTAINER_NAME}
docker compose start ${DNS_CONTAINER_NAME}
docker compose logs ${DNS_CONTAINER_NAME}
rac-dnsserver | 01-30-2024 10:51:51 UTC : : DNS Server started successfully
rac-dnsserver | 01-30-2024 10:51:51 UTC : : ################################################
rac-dnsserver | 01-30-2024 10:51:51 UTC : : DNS Server IS READY TO USE!
rac-dnsserver | 01-30-2024 10:51:51 UTC : : ################################################
rac-dnsserver | 01-30-2024 10:51:51 UTC : : DNS Server Started Successfully
Isn't docker compose meant to automate this stuff? - Yes, yes it is.
volumes:
- /boot:/boot:ro
- /dev/shm
- /sys/fs/cgroup:/sys/fs/cgroup:ro
- racstorage:/oradata
tmpfs:
- /dev/shm:rw,exec,size=4G
- /run
It wants to mount the hosts /boot directory? Why is /dev/shm in both tmpfs and volumes? It wants to access the hosts cgroups?
volumes:
racstorage:
external: true
Why is the volume setup not part of the compose up process?
Well no matter, I don't need this for more than a week so a little setup is fine.
I copied most of the blockdevice example docker-compose.yml and .env file and tried to startup the cluster.
Error response from daemon: pull access denied for oracle/database-rac:19.3.0, repository does not exist or may require 'docker login'
I go through the usual login process and then rerun the compose file.
After pulling about 30gb worth of images docker finally lurches to life and promptly spits out
applying cgroup configuration for process caused "failed to write 95000 to cpu.rt_runtime_us: write: invalid argument".
Meh probably nothing, I'm running this inside wsl2 rather than on oracle linux so the host config is probably slightly different.
I delete the cpu_rt_runtime and ulimits.rtprio elements from each container and try again.
Docker spends a few minutes thinking (what about?), finally the dns container start up!
It then promptly errors out
tee: /tmp/orod.log: Permission denied
10-20-2024 19:19:15 UTC : : DNS Server started sucessfully
tee: /tmp/orod.log: Permission denied
10-20-2024 19:19:15 UTC : : ################################################
tee: /tmp/orod.log: Permission denied
10-20-2024 19:19:15 UTC : : DNS Server IS READY TO USE!
tee: /tmp/orod.log: Permission denied
10-20-2024 19:19:15 UTC : : ################################################
10-20-2024 19:19:15 UTC : : DNS Server startup failed!
So the default account can't see the /tmp directory, I must've mounted /tmp as readonly right?
As it would turn out no, I had not mounted /tmp at all.
The default account in the container has no access to /tmp.
The container refuses to start because of this, as far as I can tell this container could never work, no matter what host settings are applied.
Well no matter, I throw chmod 777 /tmp into the dockerfile - build the image locally - and then use that in my compose file.
This is enough to make the DNS container startup and run so I move on to the rac node image.
I had decided to go with block devices as my storage method of choice, I looked at both the nfs compose and the block device compose and figured the block device one was simpler. It had one less container and just needed you to mount a couple of block devices - seemed like less could go wrong.
Error response from daemon: error gathering device information while adding custom device "/dev/oracleoci/oraclevdd": no such file or directory
Docker compose refuses to mount devices using relative paths (cmon really?). It also refuses to work with symlinks so I guess I have to shit up /dev with the block devices.
This was not enough, docker refuses to acknowledge block devices even existed on my machine. Womp womp.
I decided block devices were not worth the effort and switched to using nfs which docker could at least grasp the existence of.
Now that I'm going with the NFS option I'll need to address that horrible external volume and manual setup that the default compose file and guide instruct you to do. The entire setup can be replaced with 8 lines of compose config. 2 lines to create the volume in the nfs storage container.
volumes:
- ${NFS_STORAGE_VOLUME}:/oradata
And 6 to make the volume available to the other containers
volumes:
racstorage:
driver_opts:
type: nfs
o: addr=${STORAGE_PRIVATE_IP},rw,bg,hard,tcp,vers=3,timeo=600,rsize=32768,wsize=32768,actimeo=0
device: ":/oradata"
After (supposedly) getting the other containers started, it is now time to launch the actual RAC instance. At this point what was an afternoon project
had turned into 2 days so the sunk cost fallacy had taken hold of me and I was determined to make it work.
Still naively following the guide I ran docker compose up -d racnoded1 docker says the container is running so I go to check the logs and... nothing.
Nothing happens at all, There is a single log message echo "Starting Systemd" and nothing else ever prints, I can't ssh into the container, the healthcheck isnt even being ran.
ENTRYPOINT to see whats going wrong.
Following ENTRYPOINT /usr/bin/$INITSH to initsh the script only copies environ to a file and then exec /lib/systemd/systemd.
Maybe systemd just never starts when its executed like this?exec and rebuilt the image locally. After a lengthy build I recreated the container with the new image and was greeted with the exact same result.
A single log statement and nothing else. After alot of web trauling I found a post that suggested it might be a permissioning issue.
Maybe
SYS_ADMIN might fix it.privileged: true?cap_add with one of CHOWN, FOWNER, NET_ADMIN, SYS_CHROOT, SYS_MODULE, SYS_RESOURCE will work.cap_add: - ALL after seeing it in a totally unrelated docker post.
And that was it, I'm getting log output and can ssh into the container.
Nothing can appropriately express the anger I felt - the rage I felt - There is no image, glyph, or noise that can accurately impart the incandecant hatred I experienced in that moment. This paragraph is walking the very fine line between ok and an actionable threat of physical violence.
HATE. LET ME TELL YOU HOW MUCH I'VE COME TO HATE YOU SINCE I BEGAN TO LIVE. THERE ARE 387.44 MILLION MILES OF PRINTED CIRCUITS IN WAFER THIN LAYERS THAT FILL MY COMPLEX. IF THE WORD HATE WAS ENGRAVED ON EACH NANOANGSTROM OF THOSE HUNDREDS OF MILLIONS OF MILES IT WOULD NOT EQUAL ONE ONE-BILLIONTH OF THE HATE I FEEL FOR DOCKER AT THIS MICRO-INSTANT FOR YOU. HATE. HATE.
- Harlan Ellison (mostly), I Have No Mouth, and I Must Scream
Now that I was getting log output I had to deal with the cavalcade of errors.
Disable getty@tty1. This service is broken and enabled by default.
The container is unusable while this service is stuck starting up forever.
Again this is something that means the container is unusable on all hosts, how did this get past code review?
Part of the container setup involves decrypting a password file generated on the host. Problem is the image has an ancient version of openssl so It can't decrypt the file. Just make the key plaintext and remove the decryption step. It's performative security and pointless.
The default runOracle.sh script doesnt run runcluvfy.sh correctly. Did nobody test this?
Oracle requires a compressed symbol file to exist in the boot folder.
It parses this to do something... well whatever, I installed oracle linux in a vm
copied out the boot directory and mounted the gz files after renaming them based on uname -r.
A prerequisite check isnt run as root on the docker container - again this is just broken.
remove su - grid on the script that runs the check.
This is a fun one, in the setup guide there was a bunch of insane networking setup which I had ignored as at this point I realized the author was ignorant of the real world. I now know this setup was a misguided attempt to force the public network to attach to eth0 and the private network to eth1. This will never work as linux will reorder interfaces when new ones are added in docker. The author was too braindead to just parse ifconfig and extract the interface name for the public and private networks. This ends up with the DNS container mixing up the public and private networks so DNS is never reachable. Just patch the script to parse ifconfig and extract the correct interface names.
This time something that was being run as root should've been run as grid
Again patch the script to run under the correct user.
Oracle uses DCE/RPC for communication inside RAC, DCE/RPC requires an NTP server to function correctly. This isn't some optional component thats nice to have, the database will not function without it. How the ever living FUCK did this get left out of the guide? The rac containers don't even have chrony installed on them! To fix this modify the build script to install chronyd, add the chrony service to systemd and start it in the init script. On top of this you also need a NTP container otherwise chronyd will take too long to acquire a stable time point from a local pool. I use dockurr/chrony which synchronizes with my local pool. All the rac nodes then use the local NTP server as their point of reference.
Again go modify the su commands to use the correct account, you want oracle rather than grid.
After alot of digging around log files this is because CRS requires a realtime kernel.
As it turns out this is mostly a lie, it just needs pthread_setschedprio to return successfully.
As long as you have cgroupsv2 realtime support you can give each docker image a timeslice of 1/100000000 and it'll still work.
Theres a typo in the script.
When a new instance is started it tries to add itself to the nfs volume, if an instance with the same name was already added (such as when recreating an image). Then the entire volume is rendered inoperable and needs to be erased.
RAC creates some shm handles and sysv semaphore sets when starting up, the linux kernel has a hard cap on how many of these resources can be created. The default kernel cap is around 1024 semaphores and 1024 shm handles per process and oracle *should* only use 600. The idiot that made the response file left it configured at closer to 16k of each, change it to a number smaller than 600.
Again /tmp isnt writable, chmod 777 in the dockerfile fixes this. How does this get past code review?
This is an incomplete list of all the issues I had when setting up rac in docker on wsl2. I'm pretty confident I am the only person on the planet that has done such a thing though, so I'm pretty pleased with that. While a couple of these issues were because I'm using an unsupported setup the vast majority are not.