Has anyone experimented with those little art manikins for poses and such? I know that it isn't that hard to find reference photos for poses, but I think that the fine tuned control could be useful when you have a very particular vision in mind.
I have not! But it is an interesting idea. I know there are also these "anatomy reference" websites for visual artists, which are basically just 10s of thousands of renderings of various body types in every pose imaginable, from every angle imaginable. So that might be useful here. Another possibility I've entertained (because I happen to have one of those digital drawing tablets) would be to do quick little stick doodles showing the exact poses/scene layout I want, along with a generic character reference, and perhaps a generic reference for the look of the room, and feed it all into the AI with instructions similar to yours. This would be particularly useful if a person actually did have a "directorial vision" in mind, rather than just letting the AI come up with whatever kind of shot it wanted. And perhaps it would help with pesky little details that AIs seem to consistently mess up even with plenty of hand holding (for instance, getting a character to look at the exact right point on a hypnotic object - apparently this is hard for them to get 100% right!)
The other thing I wonder here, is that perhaps these methods would be a sneaky way to get around certain content filters. For instance with your example, just asking for a picture of an anime girl groping her own boobs might throw some AIs into a fit, but showing a posed manikin & saying "just make it like this" might get through fine