Running Oracle RAC inside Docker Compose

For a school project I needed to run an Oracle cluster with ssh access to hosts.

I had a few options for running instances.

Option 1 was out as I needed to be able to ssh into the hosts.
Option 2 was also out after I saw compute prices and found out that RAC is billed at 1$ CAD PER CPU HOUR!
Option 3 seemed promising, but after looking at the insane overhead of running VMs on a windows host I decided against it.
Option 4 would get me flogged by the rest of my family so perhaps best I didn't choose that :^)
Option 5 looks to be the only solution that wont get me bankrupted, sued, or kicked out of my house so docker it is.

The prelude

I am not the biggest fan of docker, using docker is like removing a molehill with a hand grenade. It's existance is indicative of the quality of the worlds software ecosystem as a whole, that people are willing to suffer large memory overheads and awful configuration processes; All because software is so awful that fixing it is no longer considered a viable option. Docker represents the rot that has overtaken the industry in the last decade.

That being said it is quite useful for deploying databases, so lets deploy Oracle on it. Oracle has a really good single instance docker image database/free which took me less than 10 minutes to get running with when I wanted to test my OCI (Oracle Call Interface not Oracle Cloud Infrastructure) db abstraction. Surely the RAC equivalent wouldn't be much harder to deploy, they even have a prebuilt image, a guide, and its open source on GitHub. I should have this done in an afternoon.

First Contact

I opened the RAC docker compose guide and felt my heart sink a little.

            
#---------Bring up DNS------------
docker compose up -d ${DNS_CONTAINER_NAME} && docker compose stop ${DNS_CONTAINER_NAME}
docker network disconnect ${PUBLIC_NETWORK_NAME} ${DNS_CONTAINER_NAME}
docker network disconnect ${PRIVATE_NETWORK_NAME} ${DNS_CONTAINER_NAME}
docker network connect ${PUBLIC_NETWORK_NAME} --ip ${DNS_PUBLIC_IP} ${DNS_CONTAINER_NAME}
docker network connect ${PRIVATE_NETWORK_NAME} --ip ${DNS_PRIVATE_IP} ${DNS_CONTAINER_NAME}
docker compose start ${DNS_CONTAINER_NAME}
docker compose logs ${DNS_CONTAINER_NAME}

rac-dnsserver  | 01-30-2024 10:51:51 UTC :  : DNS Server started successfully
rac-dnsserver  | 01-30-2024 10:51:51 UTC :  : ################################################
rac-dnsserver  | 01-30-2024 10:51:51 UTC :  :  DNS Server IS READY TO USE!
rac-dnsserver  | 01-30-2024 10:51:51 UTC :  : ################################################
rac-dnsserver  | 01-30-2024 10:51:51 UTC :  : DNS Server Started Successfully
            
        
Isn't docker compose meant to automate this stuff? - Yes, yes it is.
            
volumes:
  - /boot:/boot:ro
  - /dev/shm
  - /sys/fs/cgroup:/sys/fs/cgroup:ro
  - racstorage:/oradata
tmpfs:
  - /dev/shm:rw,exec,size=4G
  - /run
            

        
It wants to mount the hosts /boot directory? Why is /dev/shm in both tmpfs and volumes? It wants to access the hosts cgroups?
            
volumes:
  racstorage:
    external: true
            

        

Why is the volume setup not part of the compose up process?

Well no matter, I don't need this for more than a week so a little setup is fine. I copied most of the blockdevice example docker-compose.yml and .env file and tried to startup the cluster.

                
Error response from daemon: pull access denied for oracle/database-rac:19.3.0, repository does not exist or may require 'docker login'
                
            
I go through the usual login process and then rerun the compose file. After pulling about 30gb worth of images docker finally lurches to life and promptly spits out applying cgroup configuration for process caused "failed to write 95000 to cpu.rt_runtime_us: write: invalid argument". Meh probably nothing, I'm running this inside wsl2 rather than on oracle linux so the host config is probably slightly different. I delete the cpu_rt_runtime and ulimits.rtprio elements from each container and try again. Docker spends a few minutes thinking (what about?), finally the dns container start up! It then promptly errors out
                
tee: /tmp/orod.log: Permission denied
10-20-2024 19:19:15 UTC :  : DNS Server started sucessfully
tee: /tmp/orod.log: Permission denied
10-20-2024 19:19:15 UTC :  : ################################################
tee: /tmp/orod.log: Permission denied
10-20-2024 19:19:15 UTC :  :  DNS Server IS READY TO USE!
tee: /tmp/orod.log: Permission denied
10-20-2024 19:19:15 UTC :  : ################################################
10-20-2024 19:19:15 UTC : : DNS Server startup failed!
                
            

Fixing the DNS

So the default account can't see the /tmp directory, I must've mounted /tmp as readonly right? As it would turn out no, I had not mounted /tmp at all.

The default account in the container has no access to /tmp. The container refuses to start because of this, as far as I can tell this container could never work, no matter what host settings are applied. Well no matter, I throw chmod 777 /tmp into the dockerfile - build the image locally - and then use that in my compose file. This is enough to make the DNS container startup and run so I move on to the rac node image.

Block devices and Docker - Because it can always be worse!

I had decided to go with block devices as my storage method of choice, I looked at both the nfs compose and the block device compose and figured the block device one was simpler. It had one less container and just needed you to mount a couple of block devices - seemed like less could go wrong.

                
Error response from daemon: error gathering device information while adding custom device "/dev/oracleoci/oraclevdd": no such file or directory
                
            
Docker compose refuses to mount devices using relative paths (cmon really?). It also refuses to work with symlinks so I guess I have to shit up /dev with the block devices. This was not enough, docker refuses to acknowledge block devices even existed on my machine. Womp womp. I decided block devices were not worth the effort and switched to using nfs which docker could at least grasp the existence of.

NFS - Need For Storage

Now that I'm going with the NFS option I'll need to address that horrible external volume and manual setup that the default compose file and guide instruct you to do. The entire setup can be replaced with 8 lines of compose config. 2 lines to create the volume in the nfs storage container.

                
volumes:
  - ${NFS_STORAGE_VOLUME}:/oradata
                
            
And 6 to make the volume available to the other containers
                
volumes:
  racstorage:
    driver_opts:
      type: nfs
      o: addr=${STORAGE_PRIVATE_IP},rw,bg,hard,tcp,vers=3,timeo=600,rsize=32768,wsize=32768,actimeo=0
      device: ":/oradata"
                
            

systemd-oesn't-work

After (supposedly) getting the other containers started, it is now time to launch the actual RAC instance. At this point what was an afternoon project had turned into 2 days so the sunk cost fallacy had taken hold of me and I was determined to make it work. Still naively following the guide I ran docker compose up -d racnoded1 docker says the container is running so I go to check the logs and... nothing. Nothing happens at all, There is a single log message echo "Starting Systemd" and nothing else ever prints, I can't ssh into the container, the healthcheck isnt even being ran.

  1. I check the dockerd daemon - nothing at all related to the container.
  2. I check via vscodes docker manager - as far as vscode is concerned the container is running just fine.
  3. Maybe theres disk activity on the NFS mount and its just waiting for an initial setup? Zero bytes read/write on all volumes...
Well shit then, no logs, no terminal, no way to run diagnostics. So I start from the begining and look at the Dockerfiles ENTRYPOINT to see whats going wrong. Following ENTRYPOINT /usr/bin/$INITSH to initsh the script only copies environ to a file and then exec /lib/systemd/systemd. Maybe systemd just never starts when its executed like this?
I removed the exec and rebuilt the image locally. After a lengthy build I recreated the container with the new image and was greeted with the exact same result. A single log statement and nothing else. After alot of web trauling I found a post that suggested it might be a permissioning issue. Maybe Nothing works, the symptom doesnt even change at all, 14 hours of staring at the same log message with zero progress. I am reaching my wits end at this point - no matter what I do nothing changes. It has been 4 days and I'm just about ready to throw bricks through dockers office windows for creating such an effective torture device. As a final attempt I try cap_add: - ALL after seeing it in a totally unrelated docker post. And that was it, I'm getting log output and can ssh into the container.

This capability has NO DOCUMENTED BEHAVIOUR ON ANY DOCKER SITE

Nothing can appropriately express the anger I felt - the rage I felt - There is no image, glyph, or noise that can accurately impart the incandecant hatred I experienced in that moment. This paragraph is walking the very fine line between ok and an actionable threat of physical violence.

HATE. LET ME TELL YOU HOW MUCH I'VE COME TO HATE YOU SINCE I BEGAN TO LIVE. THERE ARE 387.44 MILLION MILES OF PRINTED CIRCUITS IN WAFER THIN LAYERS THAT FILL MY COMPLEX. IF THE WORD HATE WAS ENGRAVED ON EACH NANOANGSTROM OF THOSE HUNDREDS OF MILLIONS OF MILES IT WOULD NOT EQUAL ONE ONE-BILLIONTH OF THE HATE I FEEL FOR DOCKER AT THIS MICRO-INSTANT FOR YOU. HATE. HATE.

- Harlan Ellison (mostly), I Have No Mouth, and I Must Scream

A Sisyphean Task

Now that I was getting log output I had to deal with the cavalcade of errors.

  1. systemd is stuck starting forever or CLSRSC-00358.

    Disable getty@tty1. This service is broken and enabled by default. The container is unusable while this service is stuck starting up forever. Again this is something that means the container is unusable on all hosts, how did this get past code review?

  2. EVP_DecryptFInal_ex: bad decrypt

    Part of the container setup involves decrypting a password file generated on the host. Problem is the image has an ancient version of openssl so It can't decrypt the file. Just make the key plaintext and remove the decryption step. It's performative security and pointless.

  3. PRVG-10467 : The default Oracle Inventory group could not be determined.

    The default runOracle.sh script doesnt run runcluvfy.sh correctly. Did nobody test this?

  4. PRVE-10073 : Required /boot data is not available on node.

    Oracle requires a compressed symbol file to exist in the boot folder. It parses this to do something... well whatever, I installed oracle linux in a vm copied out the boot directory and mounted the gz files after renaming them based on uname -r.

  5. PRVG-11250

    A prerequisite check isnt run as root on the docker container - again this is just broken. remove su - grid on the script that runs the check.

  6. PRVG-11366 & PRVG-11367

    DNS 2 - Electric Boogaloo

    This is a fun one, in the setup guide there was a bunch of insane networking setup which I had ignored as at this point I realized the author was ignorant of the real world. I now know this setup was a misguided attempt to force the public network to attach to eth0 and the private network to eth1. This will never work as linux will reorder interfaces when new ones are added in docker. The author was too braindead to just parse ifconfig and extract the interface name for the public and private networks. This ends up with the DNS container mixing up the public and private networks so DNS is never reachable. Just patch the script to parse ifconfig and extract the correct interface names.

  7. PRVF-00001: Could not retrieve static nodelist. Verification cannot proceed.

    This time something that was being run as root should've been run as grid Again patch the script to run under the correct user.

  8. PRVG-1017 : NTP configuration file "/etc/chrony.conf" is present on nodes. The NTP server is not an optional component.

    Oracle uses DCE/RPC for communication inside RAC, DCE/RPC requires an NTP server to function correctly. This isn't some optional component thats nice to have, the database will not function without it. How the ever living FUCK did this get left out of the guide? The rac containers don't even have chrony installed on them! To fix this modify the build script to install chronyd, add the chrony service to systemd and start it in the init script. On top of this you also need a NTP container otherwise chronyd will take too long to acquire a stable time point from a local pool. I use dockurr/chrony which synchronizes with my local pool. All the rac nodes then use the local NTP server as their point of reference.

  9. PRKH-01014: Current user "{0}" is not the oracle owner user "{1}" of oracle home "{2}"

    Again go modify the su commands to use the correct account, you want oracle rather than grid.

  10. CRS-4535: Cannot communicate with Cluster Ready Services.

    After alot of digging around log files this is because CRS requires a realtime kernel. As it turns out this is mostly a lie, it just needs pthread_setschedprio to return successfully. As long as you have cgroupsv2 realtime support you can give each docker image a timeslice of 1/100000000 and it'll still work.

  11. PRCD-1120 : The resource for database could not be found.

    Theres a typo in the script.

  12. ORA-15045: ASM file name '+DATA' is not in reference form.

    When a new instance is started it tries to add itself to the nfs volume, if an instance with the same name was already added (such as when recreating an image). Then the entire volume is rendered inoperable and needs to be erased.

  13. CRS-0223 Resource 'ora.racnoded1.ons' has placement error.

    RAC creates some shm handles and sysv semaphore sets when starting up, the linux kernel has a hard cap on how many of these resources can be created. The default kernel cap is around 1024 semaphores and 1024 shm handles per process and oracle *should* only use 600. The idiot that made the response file left it configured at closer to 16k of each, change it to a number smaller than 600.

The Connection Manager

Again /tmp isnt writable, chmod 777 in the dockerfile fixes this. How does this get past code review?

The suffering is over

This is an incomplete list of all the issues I had when setting up rac in docker on wsl2. I'm pretty confident I am the only person on the planet that has done such a thing though, so I'm pretty pleased with that. While a couple of these issues were because I'm using an unsupported setup the vast majority are not.

Credits