Saturday, January 3, 2009

ESX + iscsi + ZFS. Provisioning options

With my ESX server(s) up and running with acceptable performance using an iscsi shared zvol, it's time to think about how I'd like to actually best provision my storage for the ESX hosts.

Which model?

When setting up FC storage all VMware admins face the architecture decision of whether to go with 1 LUN per VM or a single large LUN for all [or at least large groups of] VMs.
1 LUN per VM requires a lot of extra provisioning each time you wish to add a VM or to grow a disk, but on the up side you don't have to worry about scsi lock contension as you're effectively not using VMFS as an active/active clustered file system at all. There are possibly some performance benefits to using 1 Lun/VM if you're hitting the storage very hard as you have a dedicated queue depth per LUN too.

In my most recent deployment we opted for 1 shared LUN per RAID (2 in total) and we never looked back. Life was simpler and we didn't experience any service impacting slow downs with snapshots or other lock related activities.
Here I'll be using the same model

ZFS iscsi volume options

Within the 1 iscsi volume shared by several VMs model, what are the options that zfs gives us for this volume?
I have around 250GB free on my zpool at the moment and there are three main options as I see it:

  1. Start small with a reasonable sized zvol (50GB?), and simply grow it as VMware needs more space.
  2. Make a large (200GB?) zvol that is unlikely to be filled in the near future.
  3. Thinly provision a very large (500GB) zvol that is unlikely to be filled.

1. A 50GB zvol, that I grow each time I need more space

Positives: Only uses around as much space as I actually need, when more is needed I can grow the zvol simply by using "zfs set volumesize=xxGB"

Negatives: ESX can't directly/nicely use the additional space. When you add to the vzol size, ESX will see additional storage beyond the end of the partition that the vmfs is running within however unlike most modern filesystem there doesn't appear to be a way to grow the vmfs partition/volume to make use of the extra space directly.
You can only add another vmfs volume/partition to the end, and then span between both vmfs volumes to make one large contiguous (virtual) volume, but that really seems pretty ugly to me. I really don't know that I feel too good about VMware taking on the role of logical volume manager, I'd rather keep that intelligence on the Solaris end where I have confidence in it.

If I add another 10GB and another VMFS extent each time I want to try another VM, I can easily see my vmfs spanning many partitions very quickly all within the one zvol.... this just seems stupid and the spanning extent feature feels to me more of a last resort to work around storage that can't expand volumes on the fly.

2. A single 200GB zvol

Positives: I quickly create this once and I'll probably not need to worry about it for some time to come, with no mucky multi-extent volumes.

Negatives: I'm immediately kissing goodbye to 200GB of space on the file server, whether it's used or not. Initially I'm only looking to use maybe 3-4 VMs, each with disks of maybe 10 GB each. I have 40GB but it costs me 200GB... really not optimal! 

3. A sparse 500GB zvol

Positives: PLENTY of space for VM testing, more then I currently have available in my zpool in fact, but ZFS CAN DO!
By thinly provisioning I don't immediately write off disk space on the zpool, but ESX can format and start using what it sees as a full featured 500GB lun. ZFS uses it's COW technology so that only blocks containing newly written data are ACTUALLY written to the zpool so my sparsely provisioned xxxGB volume only uses as much data as has been written to the LUN.

Negatives: This is almost perfect except for the fact that ESX thickly provisions it's disks. This means that when I create a 40Gb VM, and install a 5GB OS, 40Gb is marked as used within zvol. ZFS doesn't know that the blocks are actually "empty" 2 levels of virtualisation up the stack. An nfs mounted VM is thinly provisioned by default, so I'm not quite getting that level of optimum volume utilisation, but NFS isn't an option... I'm just still a bit bitter about that :p

Option 3, is by far the most efficient choice of the 3 options I've evaluated, and so this is what I chose to use:

root@supernova ~]#zpool list Z
NAME SIZE USED AVAIL CAP HEALTH ALTROOT Z 2.17T 1.92T 262G 88% ONLINE -
root@supernova ~]#zfs create -s -V 500G Z/esxiscsi
root@supernova ~]#zpool list Z NAME SIZE USED AVAIL CAP HEALTH ALTROOT Z 2.17T 1.92T 262G 88% ONLINE -
root@supernova ~]#zfs set shareiscsi=on Z/esxiscsi

And...we're done!

Thursday, January 1, 2009

VMware ESX NFS performance on openstorage

There are plenty of ESX discussions about the performance of NFS vs iscsi on the web, believe me I spent a lot of time reading them all to try and get a better handle on what I was seeing here.

To summarise the general feeling over all the articles out there; the performance of NFS is much the same or slightly better than iscsi (and in most cases FC wins overall).
This is pretty much in-line with what I would have expected. Iscsi carries quite a lot of overhead and is a relatively new protocol with little in the way of speed optimisations present.
NFS has been around for decades and it doesn't have much overhead at all, with many implementations measuring near wirespeed (fishworks for example).

My environment here consists of an IBM x226 server (3Ghz, 1GB RAM, onboard Bbroadcom 1Gb NIC) running solaris nevada build 95, with 8x 300GB SATA2 disks connected to a Supermicro 8-Port SATA2 PCI-X card (AOC-SAT2-MV8).
This SATA card is a fairly cheap, dumb sata board that simply provides connectivity to the 8 SATA disks. It's important to understand that it's NOT a raid card and there is no on board cache; think of it merely as an addition 8 onboard SATA ports that the motherboard can see. ZFS makes ordinary storage like this...awesome, and in many cases faster than a hardware implementation. Many, but not all, as I discovered in testing ESX nfs performance.


The storage is laid out as follows:
root@supernova ~]#zpool status Z
  pool: Z
 state: ONLINE
 scrub: none requested
config:

 NAME STATE READ WRITE CKSUM
 Z ONLINE 0 0 0
  raidz1 ONLINE 0 0 0
  c0t0d0 ONLINE 0 0 0
  c0t1d0 ONLINE 0 0 0
  c0t2d0 ONLINE 0 0 0
  c0t3d0 ONLINE 0 0 0
  c0t4d0 ONLINE 0 0 0
  c0t5d0 ONLINE 0 0 0
  c0t6d0 ONLINE 0 0 0
  c0t7d0 ONLINE 0 0 0

errors: No known data errors
root@supernova ~]#zpool list Z
NAME SIZE USED AVAIL CAP HEALTH ALTROOT
Z 2.17T 1.91T 268G 87% ONLINE -

For those of you that haven't seen the light yet and moved to zfs and have no idea what that means, it's basically like an 8 disk Raid5.

I knew from the beginning that the storage performance wasn't going to be out of this world, but this is just a home network and it should be "good enough" for what I'm wanting to work with.

I created a filesystem called Z/VMs, and this is what it looks like after putting a few VMs on there:

root@supernova ~]#zfs get used,compress,sharenfs,shareiscsi Z/VMs
NAME PROPERTY VALUE SOURCE
Z/VMs used 27.3G -
Z/VMs compression off default
Z/VMs sharenfs on inherited from Z
Z/VMs shareiscsi off default

Results
This blog entry is being done after the fact and I didn't bother recording the exact figures as I went as I wasn't intending for this to turn into the big investigation that it turned out to be.

I immediately noticed that read performance across nfs was good, around 40MB/sec for sequential reads which is about all I tend to see from supernova to my desktop too. This was inline with what I was expecting.
Write performance though was appauling. 4-5MB/sec, sometimes I'd see 6MB/sec tops. It really was awful! From my desktop I can easily do 30+MB/sec writes to supernova, so the bottleneck wasn't network or disk throughput on the file server, so why was ESXi having such a hard time with it?

Troubleshooting this from the ESXi server's end quickly proved near impossible. Nfs doesn't show up under esxtop as disk activity at all, it's all just counted as network traffic which doesn't give many clues. In addition there are pretty much no useful knobs to tune when it comes to NFS, in turns of setting/viewing nfs mount options/flags.

Ofter nfs performance issues are down to using a buffer size that is too small or it's a protocol (udp vs tcp) issue.

Since ESXi won't tell you anything, even from the service console I was forced to watch everything from the nfs server's end, but fortunately I'm using solaris which has some great obversability tools.

nfsstat quickly showed me that ESXi was using nfs3 over TCP. Nfs4 would have been nice, but this shouldn't be the problem. Iostat wasn't hinting at and disk bottlenecks and the cpu was just ticking over.
I ran snoop to watch the nfs traffic while doing some large writes with iometer in a nfs mounted WinXP vm on the ESXi server.

[root@supernova ~]#snoop host esx1 and rpc nfs
Using device bge0 (promiscuous mode) esx1.griffous.net -> supernova NFS C WRITE3 FH=9A62 at 1917214208 for 65536 (FSYNC) esx1.griffous.net -> supernova NFS C WRITE3 FH=9A62 at 61346304 for 3072 (FSYNC) esx1.griffous.net -> supernova NFS C WRITE3 FH=9A62 at 117681664 for 1024 (FSYNC) supernova -> esx1.griffous.net NFS R WRITE3 OK 65536 (FSYNC) esx1.griffous.net -> supernova NFS C WRITE3 FH=9A62 at 1918262784 for 65536 (FSYNC) supernova -> esx1.griffous.net NFS R WRITE3 OK 3072 (FSYNC) supernova -> esx1.griffous.net NFS R WRITE3 OK 1024 (FSYNC) esx1.griffous.net -> supernova NFS C WRITE3 FH=9A62 at 117682688 for 4096 (FSYNC) supernova -> esx1.griffous.net NFS R WRITE3 OK 65536 (FSYNC) supernova -> esx1.griffous.net NFS R WRITE3 OK 4096 (FSYNC)

2 things are interesting to observe.

  1. It's using 64k (65536) byte packets, so the window size is at the correct maximum
  2. All write(3) operations are using FSYNC.

FSYNC

This was the real source of the "problem". I spent a lot of time googling this and come across a number of websites talking about zfs+nfs and zfs+databases with FSYNC, and the performance issues that come with it. This turned into a big exploration of the way writes are handled at the various layers by the difference services involved, in particular the ZFS ZIL.

At the bottom of the stack is zfs and it's disks. By default zfs caches up writes in it's journal the Zfs intent log (ZIL) and it will flush these writes to disk every 5 seconds in an aggregated/optimised write. The Zil flush operation is a O_DSYNC write, which waits until the disks themselves have confirmed that the write has completed before returning (remember that disks themselves also have caches).
This is a time consuming operation, but if it's critical that your data makes it to disk before the application moves on (think databases), then it's entirely appropriate.

In a RaidZ, a write operation occurs across several disks and the entire strip has to be read to calculate the parity for the stripe before updating it. Raid-5 write performance is always going to be less than optimal because of this. RaidZ has a variable length stripe which helps a lot, but the bottom line is that with a single width RaidZ you'll only get the IOPs of a single disk.

Putting this all together on a system with no DRAM based disk write cache as I have here, and a O_DSYNC/FSYNC operation is going to give the net write performance of a single unbuffered disk or less. Fortunately most applications don't use FSYNC.

NFS writes can optionally have the fsync bit set on write operations, which works it's way down the layers to the filesystem (zfs), it's journal (zil), and then the disks/array itself (my sata disks).
When an NFS client requests a fsync write the nfs server cannot confirm the completion of the write request until that data has actually been confirmed as written on disk and zfs honours this.

Based on my snoop results, *all* ESX writes are using fsync which mean that for every IO write request in a VM the nfs server has to flush all writes to disk before continuing with the next request. Installing an OS for example will issues hundreds of write fsync requests per second which is just annihilating the write throughput.

Given the critical nature of the data running in virtual machines, I don't think ESXi is doing anything wrong here. Data being lost midtransaction could have disasterous results for the VMs further up the stack and the only way for ESXi to gaurantee data integrity is to use fsync for all writes.

This all makes sense.... so why is everyone else reporting that nfs performance is on par with iscsi when from what I've seen here, it's a disaster due to fsyncs! Funnily enough my googling turned up a few other people asking the same question, also reporting around 4-5MB/sec on writes, with many talking about using linux nfs storage.
Unfortunately no one had replied to these threads to I had to do a bit more head scratching to get the answer.

I think the answer is this; I'm doing nfs against a commodity server, with commodity disks which means that my writes are done using write-through while most ESX(i) installations will be against commercial SAN/NAS such as netapp appliances, which will operate in write-back mode as they have a battery backed cache of DRAM of NVRAM that can survive a power outage without data loss.

Write Caches

In write-through mode a FSYNC write is written to disks before returning with a successful IO, which is just exasserbated over a higher than local latency network storage system such as NFS.
In write-back mode, as soon as the data is in DRAM or in NVRAM the array will return a successful IO allowing the client keep hammering those IOs through even though they haven't actually made it to disk yet. This is actually transactionally safe because the battery backup will ensure that those writes that didn't actually make it to disk are replayed from cache (still live due to the battery) onto the disks as soon as power returns.
Naturally write-back caching makes a huge difference to write performance latency.

My home system is using simple disks with no additional write caching so I can't do write-back caching.....or can I?

ZFS is very configurable, and there is a "knob" that you can change on the fly to disable flushing fsync transaction to disks. By disabling the ZIL, zfs will effectively ignore fsync/O_DSYNC requests, or put another way it changes your zpool to behave like writeback storage. Now, this is a very BAD idea and should never be used in production as it will cause corruption for nfs clients in the event of a power outage. Don't do it! Really, don't. More information can be found here.

I wanted to confirm that my understanding of all this was in fact accurate so while doing a large file copy within the VM as root I issued "echo zil_disable/W0t1 | mdb -kw", which disables the ZIL globally. Straight away, mid copy the file copy performance rocketed up and I started seeing more like 25-35MB/second. Woohoo, so it's the lack of a write-back cache that's killing nfs performance in my environment.

Obviously leaving thi ZIL disabled isn't a safe thing to do and I don't want corruption so I put it back again with "echo zil_disable/W0t0 | mdb -kw".
It proved that I was correct in my assumption though, it was the write-through behaviour of zfs + nfs that was killing my performance.

ZFS does have the ability to put the zil onto alternate storage while keeping your data on the main zpool. Putting your zil onto a battery backed RAM device or a solid state disk will do wonders for this kind of loading so I could likely solved this problem by putting my zil on a seperate SSD.
Right at the moment there is a rather major bug with this functionality in zfs; once a zil (slog) has been setup on an alternate vdev it can't be removed again. I'll be keeping an eye on that bug, and once fixed I'll seriously look into going ahead with doing this.

Having the zil on very low latency storage makes the most difference for O_DSYNC/fsync operations but it will benefit a large range of fs loads.

What next? Iscsi testing

Having proved that transactionally stable nfs storage on supernova is going to be painfully slow until I add additional hardware I turned to my other option for storage on supernova, iscsi.

The short of it is that iscsi is performing perfectly on both reads and writes which I'm very pleased to see. I haven't yet had a big dig into what operations are going on with iscsi in terms of disk flushes on write, but whatever the differences are the user facing result is much better performance.
I still think nfs is far more flexible and would offer many great advantages for management but until I can get the performance on par with iscsi, I'll stick with iscsi.

I think the best news is that opensolaris still does remain a great platform for use with ESX. With project comstar even FC is an option, though I don't have the hardware here to play with that :P

Whitebox burn in and VMware

With everything assembled I went straight into the benchmarking and testing.
The bios in on the XFX is very nice to use, which heaps of information and very specific voltage/FSB adjustments.

I quickly discovered that I could take my 2.83 up to 4GHz, I seem to remember getting to about 4.08 before it would stop posting and I didn't want to push to the voltage on the cores too highly.

RAM

From here I went down on to discover a number of very interesting points about burn in testing and memory speed. I purchased brand name DDR1066MHz RAM and with the motherboard supporting all the way up to DDR 1200MHz RAM, I should be able to run it at it's "factory" DDR1066MHz settings, right?...

It turns out that corsair advise running the RAM at a whopping 2.1V (1.8V is the default). I wouldn't have dreamed of pushing my RAM's voltage that high for fear of breaking it, but sure enough I confirmed this information on their website.

Frustratingly, even at 2.1V memtest86 was still giving me RAM errors even with the CPU/FSB running at it's factory defaults. Even running at 2.15V didn't cure the problem.
I thought it might just be a bad Dimm, so pulled a pair out... no problem. I switched pairs to the presumably faulty pair...no problem there either.
Wait, what???!

Then I remembered discovering something similar back when I built my AMD desktop some years ago. The memory controller on the CPU (AMD remember) just couldn't drive all 4 Dimms at their uprated speeds so if you were to run 4 Dimms you had to drop the DDR speed down. With this system being an Intel with a dedicated off-CPU memory controller I never even gave this a thought, but here I was having the exact same problem symptoms. I'd already upped the SPP voltage, and the FSB (despite running at a stock 333Mhz)... no dice.

I did explore running at DDR 1000 but after eventually getting another crash I finally conceeded default and fell back to DDR800 with my DDR1066 RAM. Very disappointing, but I'm not sure that it's the RAM at fault specifically and RMAing this would be very challenging.

CPU limits?

On the CPU side, I had settled on 3.83Ghz, which seemed to be pretty stable with the CPU temps in the high 50s at idle. Unfortunately I couldn't easily monitor the CPU temps under load save for thrashing it and then really quickly rebooting then jumping into the BIOS and checking the cpu temperatures. There are some windows tools to monitor the BIOS lifesigns, but I was using ubuntu/ESX at all times.

After an extended period of load at 3.83Ghz I'd get weird failures, and even at 3.6Ghz I'd have problems once I'd installed the servers in my server room. I guess there isn't the same airflow in there so it's a little less forgiving. For testing here I was running memtest in a 2CPU VM with a 2CPU xp VM running prime95.
Prime95 has turned out to be extremely useful. It was forever picking up errors in it's own calculations hinting that there was a memory or CPU error; normally memory. What's interesting is that memtest wasn't picking anything up so my guess is that it's something to do with the FPUs being hammered in addition to the RAM itself, while memtest is just testing the RAM.

Either way, prime95 was frequently telling me that I had issues even when everything else appeared to be running OK. I even installed server 2008 in another VM on the same host while it was telling me there were issues. I really started to conclude that it was just prime95 that was having the issues, but sure enough a crash/purple screen would come along soon enough.

3.4Ghz has been very stable, with prime95 & memtest running in VMs all night long without issues, so I've settled on that. The voltage in the bios for the CPU has been set to 1.250V, but it's actually getting around 1.19V after what I've since discovered to be known as "voltage droop".

I think the default voltage is 1.15, so the CPUs are hardly any warmer for the extra 500Mhz/core. 2Ghz is worth having in my book!

VMware ESXi

I did quite a bit of googling to try and work out if my XFX MG-V780-ISH9 motherboard would be supported by ESX so for the benefit of any others that may hit this blog; the NICs both work using the forcedeth driver. IDE/PATA also works, though you'll have to hack the ESXi installer if you wish to install onto them as ESXi actively ignores PATA disks during install weirdly enough, while it will let you use them for vmfs stores. I found a guide somewhere to bypass the IDE restriction at install time, but I'v ultimately ended up using USB keys for booting anyway. (Less noise/heat... and maybe even faster?)

I'm not sure if the SATA ports work, I haven't tried them.

Having spent a bit more time hacking at ESXi now, I must say it's very annoying to use. We all know about the "u n s u p p o r t e d" hack to use a console, and yes you can enable SSH access (for now at least), but the service console has very little in it. Yes, I know this is kinda the whole point, but boy it makes troubleshooting stuff a pain in the ass!

Another odd ESXi specific oddity is the networking for service console. Under ESX you have 3 types of networks.
VM networks
Service console networks
VMkernel networks.

ESX merges the Service Console and the VMkernel into one.
I've been using NFS as the backing storage to get things up and running quickly and I discovered very early on that the performance metrics for the disk usage simply don't exist with NFS, it all shows up as network traffic.
If you want to know how much disk load an individual VM is generating, you can't look at it's disk performance information (there isn't even a drop down for disk), but furthermore the network metrics are just literally for the VM's actual network (not it's underlying nfs traffic).

This just leaves monitoring it at the host level, and with the service console network data mixed in with the nfs traffic it all gets a bit mixed up.

Rather a weird way of doing things VMware!

VMware Whiteboxen

I've been in denial for a while, but the time finally came to shell out some cash on faster servers for my home testlab. It currently exists as a mess of old desktops, and even older Compaq servers (before hp rebranded them).

While everything has keeping up with what I need it to do, I really want to start spending more time working with virtualisation technologies and bringing myself up to speed again with some of the later MS technology (server 2008/sql 2008/etc) and the fact is that 4-8 year old hardware just can't cut it any more.

I've been running Xen for many years now with great success but I really wanted chance to start testing other hypervisors such as VMware's, and to test some of the newer MS products in VMs. Windows VMs of course require hardware assistance, so my older hardware won't do (even slowly), regardless of the hypervisor used (with the possible exception of qemu, which isn't technically a hypervisor anyway).

Real Servers or whiteboxes again?

I had a quick look into the pricing and options for buying a cheap commercial server, from the likes of Dell having had good success with my last purchase of an entry level IBM x226 as my file server. The big killer when it comes to servers is the RAM pricing. VM hosts need lots of RAM and I wanted a minimum of 8GB per server. ECC RAM is not cheap and that's the only kind you can get even with entry level servers so having reviewed the options I ended up going down the DIY/whitebox path.

A friend from work builds PCs all the time and he was helpful enough to suggest a base config which I just tweaked a bit to come out with my final component list.

The final configuration was 2:

Intel Core 2 Quad Q9550 2.83Ghz CPU
XFX MG-N780-ISH9 motherboard
Coolermaster Elite 330 Black Case
Lite-on 20A4P PATA DVD Burner
Corsair XMS2 DDR2 1066MHz 2GB RAM pair (x2)
Silverstone Olympia OP700 700W Power Supply
ASUS 8400GS silent video card

The CPU was the fastest I could find readily, though I think there is a 3Ghz version out there somewhere.The motherboard is very much a gaming motherboard, in fact it's triple SLI. Obviously the video prowess wasn't the goal, but given that it's built for high throughput and overclocking it will mean that it's a very stable board. The powersupply is bigger than I need which should add to the reliability, and the case was pretty much the cheapest one I could find.
The video card and DVD burners will probably be repurposed and shuffled at a later time, but to get things started I needed both to initially build the system. I have plans to use one of the video cards in my HTPC for h264 decoding as covered here. I'm actually using the exact same card they used to be sure that it will work :)
I went for 8GB of RAM, which is actually the maximum that these motherboards support anyway. I also opted to buy the faster DDR1066 versions for some possible overclocking and some extra margin above DDR800.

With everything "overspeced", they should be very reliable.

I haven't included any hard drives as initially I'll be running ESXi, and attempting to use it from a USB key. I purchased a pair of "high speed" 4GB usb keys for this.

File Server RAM

Finally I also ordered another 4GB of RAM for my file server. It's been humbly chugging along with 1GB serving up nfs/cifs while also running 4 zones, but once I start throwing VM traffic at it too (iscsi/nfs) it won't have the RAM to cache anything and the performance is going to suffer big time. Being a server, I had to pay way too much for the ECC RAM. I did find some Kingston aftermarket DIMMS that are gauranteed to work rather than the stupidly priced IBM OEM RAM. The documentation on the RAM configuration for the x226 is very confusing so I still don't actually know if I'll be able to use my current pair of 512MB DIMMs in conjunction with the 2x2GB DIMMs that I've ordered. Worst case, I'll have 4GB of RAM, best case, 5GB. I can live with that.

Hardware delivery

While most of my new "server" hardware arrived before christmas, the CPUs were on backorder; especially frustrating as they WERE in stock when I placed my order so that I could work on this over the holidays.
This mean that I couldn't actually start on anything until the 29th December.
It turns out that my server RAM has been delayed too, which is annoying but I can start testing everything with only 1GB.

Sunday, October 26, 2008

Centralised Authentication on Solaris part #3

With the base software now installed, it's time to test it.

First off, we need to check if the web server is started. Is isn't:

root@ds1 bin]#/usr/sbin/smcwebserver status
Sun Java(TM) Web Console is stopped

root@ds1 bin]#/usr/sbin/smcwebserver start
Starting Sun Java(TM) Web Console Version 3.0.3 ...
The console is running

I was then able to login into the web console at https://ds1:6789 using my standard local user account. To initialise the Directory Service Control Center apparently requires root access, so against my better judgement I logged into the web interface as root.

The Control Center did it's thing, and to my delight it even recommended that I return back to using a non-privileged user again now that the initialisation was complete.

Initial indications were that everything seemed mightily slow, but it may speed up with time.
The authentication to the webconsole can be done as any local user and once you select the DSCC, you're asked for admin authentication to the DS itself. Makes sense, so basically the Sun directory server is just piggybacking on the existing web management console for remote administration. There doesn't appear to be a local ldap/management client.

It quickly became obvious that the web interface didn't have any directory servers registered. It seems that I setup what was needed for the DS server itself, but not an actual ldap directory.

I tried created a new directory, but failed with Could not contact the DSCC agent on ds1. Use the command cacaoadm to check that the DSCC agent is installed and running on port 11162.

So:

root@ds1 bin]#cacaoadm status
Cannot find property: [cacao.embedded].

Doh!
Google provided http://bugs.opensolaris.org/view_bug.do;jsessionid=4694b06edf8cd25d148e6a54c8fa?bug_id=6745235
which shows this as being a bug in snv_97. This host is using snv_95, but I seemed to be having the same issue:

root@ds1 bin]#pkginfo -l SUNWcacaort | grep VERSION
VERSION: 2.0,REV=15

and yet the snv_95 DVD came with:
root@angelous Product]#cat /mnt/tmp/Solaris_11/Product/SUNWcacaort/pkginfo | grep VERS
VERSION=2.2.0.1,REV=2008.06.06

So I replaced SUNWcacaort, SUNWcacaosvr and SUNWcacaodtrace. There were errors along the way, but at least the command didn't return errors now.
The service isn't started by default:
root@ds1 ~]#cacaoadm enable
root@ds1 ~]#cacaoadm status
default instance is ENABLED at system startup.
default instance is not running.
root@ds1 ~]#cacaoadm start
root@ds1 ~]#cacaoadm status
default instance is ENABLED at system startup.
Smf monitoring process:
12649
12650
Uptime: 0 day(s), 0:1

Lots more RAM was being eaten up now, but moving on...

Cacoa was only listening on localhost, so to get dsee working with it, I had to rerun the registration from dsccsetup

root@ds1 ~]#cacaoadm list-params | grep network-bind
network-bind-address=127.0.0.1

root@ds1 ~]#/opt/SUNWdsee/dscc6/bin/dsccsetup status
***
DSCC Application is registered in Sun Java (TM) Web Console
***
DSCC Agent is not registered in Cacao
***
DSCC Registry has been created
Path of DSCC registry is /var/opt/SUNWdsee/dscc6/dcc/ads
Port of DSCC registry is 3998
***
root@ds1 ~]#/opt/SUNWdsee/dscc6/bin/dsccsetup cacoa-reg
Invalid subcommand cacoa-reg
For more information, see dsccsetup --help.
root@ds1 ~]#/opt/SUNWdsee/dscc6/bin/dsccsetup cacao-reg
Registering DSCC Agent in Cacao...
Checking Cacao status...
Stopping Cacao...
Enabling remote connections in Cacao ...
Starting Cacao...
DSCC agent has been successfully registered in Cacao.
root@ds1 ~]#cacaoadm list-params | grep network-bind
network-bind-address=0.0.0.0

Better!

Creating the new ldap server failed the first time, for two reasons.
1. The user that I was using didn't have the right to create a service with a low port number.
Solved using RBAC:
root@ds1 ~]#usermod -K defaultpriv=net_privaddr ds

2. The directory that I specified for the new domain didn't exist yet, and the used didn't have permission to create it. I rather just expected this to work.
root@ds1 SUNWdsee]#mkdir griffous.net
root@ds1 SUNWdsee]#chown ds:ds griffous.net/
root@ds1 SUNWdsee]#cd griffous.net/
root@ds1 griffous.net]#pwd
/opt/SUNWdsee/griffous.net
However this too didn't quite work, because the directory already existed. Finally I granted ds access to /opt/SUNWdsee, just to get the directory created.

I didn't manage to get the server started on a priviledged port, but it did come up on a non-priviledged one. I'll troubleshoot that further tomorrow.

Saturday, October 25, 2008

Centralised Authentication on Solaris part #2

Installing a whole root zones take time. A lot of it.

# time zoneadm -z ds1 install
A ZFS file system has been created for this zone.
Preparing to install zone .
Creating list of files to copy from the global zone.
Copying <210541> files to the zone.
Initializing zone product registry.
Determining zone package initialization order.
Preparing to initialize <1334> packages on the zone.
Initialized <1334> packages on zone.
Zone is initialized.
The file contains a log of the zone installation.

real 65m58.822s
user 2m38.165s
sys 8m11.322s

Given just how long it took, one starts to wonder at the advantages of using zones at all, over just using stand alone virtual machines. A full VM is much more portable, and I wonder just how much management you save by using a whole root zone.
A topic for another day...

With my newly initialised zone, I proceeded to run the installer again.
This post is going to be very text heavy, but it serves as some self documentation for me, and maybe an outline for anyone else that hits this.

Choose Software Components - Main Menu
-------------------------------
Note: "* *" indicates that the selection is disabled

[ ] 1. Web Server 7.0 Update 1
[ ] 2. Directory Preparation Tool 6.4
[ ] 3. Application Server Enterprise Edition 8.2 Patch 2
[ ] 4. Directory Server Enterprise Edition 6.2
[ ] 5. Monitoring Console 1.0 Update 1
[ ] 6. High Availability Session Store 4.4.3
[ ] 7. Access Manager 7.1
[ ] 8. Message Queue 3.7 UR2
* * Java DB 10.2.2.1
[ ] 10. All Shared Components

Enter a comma separated list of products to install, or press R to refresh
the list [] {"<" goes back, "!" exits}: 4
Enter a comma separated list of components to install (or A to install all )
[A] {"<" goes back, "!" exits}

*[X] 1. Directory Service Control Center
*[X] 2. Directory Server Command-Line Utility

Next came the share components that failed in the sparse zone:

Shared Component Upgrades Required
-----------------------------------

The shared components listed below are currently installed. They will be
upgraded for compatibility with the products you chose to install.

Component Package
--------------------
Cacao SUNWcacaort
2.2.0.1 (installed)
2.0:PATCHES:123896-03 (required)
Cacao SUNWcacaowsvr
2.2.0.1 (installed)
2.0:PATCHES:123897-03 (required)
JavaActivationFramework SUNWjaf
8.0.0.0 (installed)
8.1 (required)
JavaMail SUNWjmail
8.0.0.0 (installed)
8.1 (required)
SunWebConsole SUNWmctag
3.0.2 (installed)
3.0.2:PATCHES:125953-05 (required)
SunWebConsole SUNWmconr
3.0.2 (installed)
3.0.2:PATCHES:125951-05 (required)
SunWebConsole SUNWmcon
3.0.2 (installed)
3.0.2:PATCHES:125953-05 (required)
SunWebConsole SUNWmcos
3.0.2 (installed)
3.0.2:PATCHES:125951-05 (required)
Ant SUNWant
11.11.0 (installed)
11.12.0 (required)

Enter 1 to upgrade these shared components and 2 to cancel [1] {"<" goes
back, "!" exits}:1

I told it to upgrade everything it needed.

Installation Directories
------------------------

Enter the name of the target installation directory for each product:


Directory Server [/opt/SUNWdsee] {"<" goes back, "!" exits}:
Directory Preparation Tool [/opt/SUNWcomds] {"<" goes back, "!" exits}:


Checking System Status

Available disk space... : Checking .... OK

Memory installed... : Checking .... OK

Swap space installed... : Checking .... OK

Operating system patches... : Checking .... OK

Operating system resources... : Checking .... OK


System ready for installation


System Ready for Installation. Memory detection is disabled in a non-global zone.

Neat, it worked out that it's not real. On with the installation, nearly.

Screen for selecting Type of Configuration

1. Configure Now - Selectively override defaults or express through

2. Configure Later - Manually configure following installation


Select Type of Configuration [1] {"<" goes back, "!" exits} 1

I opted to configure everything up front, since I may not know how to change the defaults easily later

Specify Common Server Settings

Enter Host Name [ds1] {"<" goes back, "!" exits}
Enter DNS Domain Name [griffous.net] {"<" goes back, "!" exits}
Enter IP Address [192.168.1.74] {"<" goes back, "!" exits}
Enter Server admin User ID [admin] {"<" goes back, "!" exits}
Enter Admin User's Password (Password cannot be less than 8 characters) []
{"<" goes back, "!" exits}
Confirm Admin User's Password [] {"<" goes back, "!" exits}
Enter System User [ds] {"<" goes back, "!" exits} ds
Enter System Group [ds] {"<" goes back, "!" exits} ds

Directory Server: Create Directory Instance

Directory Server Console requires Directory Server, but does not require
a directory instance.

Although not a requirement, you can create a directory instance now
during installation.

Create a directory instance (in addition to installing Directory Server)?


1. Yes
2. No

Enter 1 or 2 [1] {"<" goes back, "!" exits} 1

I opted to setup the directory while at it, accepting all the defaults.

Directory Server: Specify Instance Creation Information

Enter Instance Directory [/var/opt/SUNWdsee/dsins1] {"<" goes back, "!"
exits}
Enter Instance Port [389] {"<" goes back, "!" exits}
Enter Instance SSL Port [636] {"<" goes back, "!" exits}
Directory Manager DN [cn=Directory Manager] {"<" goes back, "!" exits}
System User [root] {"<" goes back, "!" exits} ds
System Group [root] {"<" goes back, "!" exits} ds
Enter Instance password (At least 8 characters long) [] {"<" goes back, "!"
exits}
Retype Password [] {"<" goes back, "!" exits}
Enter Suffix [dc=griffous,dc=net] {"<" goes back, "!" exits}
Ready to Install
----------------
The following components will be installed.

Product: Java Enterprise System Identity Management Suite
Uninstall Location: /var/sadm/prod/SUNWident-entsys5u1
Space Required: 138.79 MB
---------------------------------------------------------
Sun Java(TM) System Directory Preparation Tool
Sun Java(TM) System Directory Server Enterprise Edition 6.2
Sun Java(TM) System Directory Server Enterprise Edition 6.2 Command-Line
Utilities
Java Enterprise System Directory Server 6.2 Core Server
Java Enterprise System Directory Service Control Center


1. Install
2. Start Over
3. Exit Installation

What would you like to do [1] {"<" goes back, "!" exits}?1

This failed, so I tried again, this time with everything as root.

Java Enterprise System Identity Management Suite
|-1%--------------25%-----------------50%-----------------75%--------------100%|


Installation Complete

Centralised Authentication on Solaris

It's a big topic. Some might even call it a bit daunting.

Most importantly this is a subject that I didn't feel I knew enough about, and so began the journey of discovery to learn what one must do to have single sign on (SSO) under Solaris.

Microsoft's AD really does make it a bit easy for windows admins - there really is just the one choice, and like it or not you're going to be doing it the one way. About the only weird thing is that they persist with making the very standard task of setting up a domain controller, a command line app (dcpromo.exe), despite windows being very obviously a GUI centric world.

On the Solaris/Linux front we're too spoilt for choice. There are a good number of centralised directory/LDAP systems out there and I spent the best part of a day just trying to catch up on all the naming conventions and versions.

The major players appear to be:
Apache Directory Server
OpenDS
OpenLDAP
Sun Java System Directory Server

I had a look at the last three in some depth.
OpenDS is an entirely java based ldap server, which is a bit...odd, but more importantly to me, it doesn't appear to have any kind of a integrated gui front end whatsoever.
Now I know what you're thinking, it's Solaris, if you don't want command line, then pick up your toys and go play with Windows!

I've already deployed OpenLDAP on linux, and have run it for around 4 years powering my home network. It's been solid but it's really painful having to use command line tools for even the most mundane of updates. By the time you throw some kind of SSL into the mix, it becomes really ugly to manage and maintain. I'm sorry - this time around I wanted a GUI and that's that.

Which brings us to the Sun Java Directory Server - a product which goes by so many many names it's VERY hard to figure out what you are working on. Being a Sun product designed for Solaris it did seem the obvious tool for the job right from the start, but I wanted to at least give some of the other choices some review before making a start.

Having made the decision to go with the Sun DS, I created a zone on supernova and started the install.

Details in the next post.