unix sysadmin archives
Donation will make us pay more time on the project:
          

Showing posts with label Solaris. Show all posts
Showing posts with label Solaris. Show all posts

Thursday, 4 July 2013

ufsrestore -if

Another helpful unix utility is the ufsrestore.
It's very handy, specially when you only have to restore a file or a specific directory.
This is preferable than netbackup for it's easiness.
But of course a ufsdump is needed for you to make use of this tool.

Here are the highlights of how to use it. Let say we want to restore /etc
#ufsrestore -ifv
add /etc
extract
Specify next volume #: 1 comment: in most cases is volume 1
set owner/mode for '.'? [yn] n comment: NO when restoring in directory other then one from which files were dumped
comment: YES if restoring in same directory from were dump was performed.

ufsrestore > quit

Thursday, 3 May 2012

How to Replace System Board for Sun Fire E6900 Systems

Applies to:
Sun Fire 4800 Server - Version: Not Applicable and later [Release: N/A and later ]
Sun Fire 4810 Server - Version: Not Applicable and later [Release: N/A and later]
Sun Fire 6800 Server - Version: Not Applicable and later [Release: N/A and later]
Sun Fire E4900 Server - Version: Not Applicable and later [Release: N/A and later]
Sun Fire 3800 Server - Version: Not Applicable and later [Release: N/A and later]
Information in this document applies to any platform.

H/W ON-SITE Action Plan #1. Parts: 540-6295
NOTE: THIS AUTOGENERATED CREATED USING https://actionplans.us.oracle.com/atr/.

***********************************************************
************ Start Hardware Onsite Action Plan ************


A. DISPATCH INSTRUCTIONS

A1. WHAT SKILLS DOES THE ENGINEER NEED (IS A SITE ENGINEER AVAILABLE?):

A2. PARTS REQUIRED: USE INFORMATION IN TASK UNLESS ENTERED BELOW
Part number: [F] 540-6295
Part location: SB4
Quantity: 1
Description: CPU/MEM W/ 4 US IV 1.35GHZ, 0GB (FRU)
Prior part DOA: No
Alternate parts: 540-6803
SPECIAL INSTRUCTIONS: Verify the new board's firmware matches that of the System Controller and other boards in the configuration.
See http://sunsolve.sun.com/search/document.do?assetkey=1-61-214805-1 for details.

A3. DELIVERY REQUIREMENT:
Preferred Onsite Time: Within Service SLA

A4. ONSITE VISIT DETAILS:
Account name:
Contact Name:
Contact Telephone #:
Email address:
Street Address:
City:
State:
Country:
Postal Code:
Alt. Contact name:
Alt. Contact email:
Alt. Contact phone:
Special instructions:

B. FIELD ENGINEER INSTRUCTIONS

NOTE : READ MANDATORY NOTES SECTION OF ACTION PLAN.
This Action Plan is not complete until all mandatory actions outlined below have been competed.

B1. PROBLEM OVERVIEW:
General problem: There is a component failure
Fault for part 540-6295: NA
*** Start System Error Message ***
cat-a:SC> showchs -c SB4 -v
Total # of records: 1
Component : /N0/SB4
Time Stamp : Sun Dec 04 19:26:26 EST 2011
New Status : Faulty
Old Status : OK
Event Code : HW
Initiator : ScApp
Message : 1.E6900.FAULT.ASIC.CHEETAH.AFSR_2_HI_ISAP.71191111.20-16.1

*** End System Error Message ***

B2. WHAT ACTIONS DOES THE ENGINEER NEED TO TAKE:
In this Document

Goal

Solution


Oracle Confidential (INTERNAL). Do not distribute to customers

Reason: FRU CAP

Applies to:
Sun Fire 4800 Server - Version: Not Applicable and later [Release: N/A and later ]
Sun Fire 4810 Server - Version: Not Applicable and later [Release: N/A and later]
Sun Fire 6800 Server - Version: Not Applicable and later [Release: N/A and later]
Sun Fire E4900 Server - Version: Not Applicable and later [Release: N/A and later]
Sun Fire 3800 Server - Version: Not Applicable and later [Release: N/A and later]
Information in this document applies to any platform.


Goal
How to Replace System Board for Sun Fire 3800, 4800, 4810, 6800, E4900, and E6900 Systems

******************************************************************************

To report errors or request improvements on this
procedure,

please go to http://support.us.oracle.com
and put a comment on Doc ID: 1306577.1

******************************************************************************

Solution
DISPATCH INSTRUCTIONS

WHAT SKILLS DOES ENGINEER NEED:

ScApp, lom

Task Complexity: 4

Time Estimate: 60 minutes

FIELD ENGINEER INSTRUCTIONS

CAP PROBLEM OVERVIEW:
System Board Failure

WHAT STATE SHOULD SYSTEM BE IN TO BE READY TO PERFORM RESOLUTION ACTIVITY?

Examples use a board location of '#'

1) See if DR can be used; If the board is listed in 'cfgadm -av | grep -i perm' output, you can't use DR.

2a) If able to DR, issue 'cfgadm -c disconnect N0.SB#'

2b) If unable, issue 'init 0' and then 'poweroff sb#' at SC prompt

WHAT ACTION DOES ENGINEER NEED TO TAKE:

You will need to move DIMMs from the 'old' board to the 'new' one (same slots).

1) Perform physical SB replacement per Service Manual

2) poweron SB# at SC prompt

3) showchs -b at SC prompt

4) Reset any 'Suspect' or 'Faulty' components to 'ok' from the Main SC or from the lom prompt:
setchs -s OK -r 'SR number' -c <comp>

NOTE: If ScApp 5.20.15 or higher, service mode access IS NO LONGER REQUIRED to execute setchs.
If < 5.20.15, contact service to obtain a Service Mode password or generate one yourself at https://modepass.us.oracle.com
(a backup server is also available from https://modepass-bak.us.oracle.com)
Repeat 'setchs' command until all components are 'ok'.

Verify 'showchs -b' is empty.

5) Verify new board firmware matches existing boards & SC(s) ('showboards -p proms' at SC prompt).

If needed, copy firmware from a like board 'flashupdate -c (source board) (destination board)'

6) Consider running extended POST (On domain issue 'eeprom diag-level=max' or at ok prompt 'setenv diag-level max').

7) If you're replacing a COD (Capacity on Demand) enabled board refer to
Sun Fire[TM] 12K/15K/E20K/E25K/F3800/Fx800/Ex900/ servers: How to replace a COD CPU/memory board (Doc ID <a href="<<INLINE_NOTE:
1002102.1
>>">
1002102.1
)
for the needed step to follow

8a) If able to DR, issue 'cfgadm -c configure N0.SB#' at domain level

8b) If unable to DR, issue 'setkeyswitch -d (domainID) off' followed by 'setkeyswitch -d (domainID) on'.

9) Monitor POST.
* If new errors are detected, collect POST and contact Support.

OBTAIN CUSTOMER ACCEPTANCE

WHAT ACTION DOES CUSTOMER NEED TO TAKE TO RETURN SYSTEM TO AN OPERATIONAL STATE:

Boot system if not already booting.

REFERENCE INFORMATION:

Replacement procedures are documented in the Service Manuals:

* Chapter 8, 3800/48x0/6800 Manual http://download.oracle.com/docs/cd/E19095-01/sf3800.srvr/805-7363-15/805-7363-15.pdf

* Chapter 8, E4900/E6900 Manual http://download.oracle.com/docs/cd/E19095-01/sfe4900.srvr/817-4120-13/817-4120-13.pdf


B3. SHOULD DYNAMIC RECONFIGURATION BE USED:

B4. IS OUTAGE REQUIRED AND AGREED TO BY THE CUSTOMER: Yes

B5. NOTICES THAT ENGINEER MUST TAKE INTO ACCOUNT:
ROHS NOTICE: This system has NOT been adequately identified during remote diagnosis for the purposes of RoHS. You must check the system's RoHS compliance by referring to information within FIN 102250, or by verifying with your support centre before commencing service.

B6. ADDITIONAL COMMENTS:

B7. WHAT TROUBLESHOOTING TESTS WERE DONE:


C. GENERAL ACTION PLAN INFORMATION
Action plan for case:
Action plan reference number: 1 (always reference with case id)
Affected product: SUN FIRE E6900
Platform version: N/A
(Please update MOS with correct serial number if necessary)


************* End Hardware Onsite Action Plan *************
***********************************************************


**************** GENERAL INSTRUCTIONS FOR THIS ACTION PLAN ***************
**************************************************************************
Make sure a new explorer [explorer -w all,interactive,scextended] is run and
submitted to proactive.central after installing the new parts. Email explorer to:
explorer-database-americas@sun.com - Americas
explorer-database-emea@sun.com - EMEA (Europe, Middle East, Africa)
explorer-database-apac@sun.com - APAC (Asia, Pacific)
**************************************************************************
***********************************************************************

Monday, 13 February 2012

When downtime is inevitable

WHAT STATE SHOULD THE SYSTEM BE IN TO BE READY TO PERFORM THE RESOLUTION ACTIVITY :
Customer should shut OS down gracefully (using "init 0").
Then, At the OK prompt Type #. (in key sequence) to get into Alom.
Run the following commands from ALOM:
1. setlocator on
2. poweroff

WHAT ACTION DOES THE ENGINEER NEED TO TAKE:
1. Shutting System Down.
2. Extending Server to Maintenance Position.
3. Performing Electrostatic Discharge Prevention Measures.
4. Disconnecting Power From Server.
5. Removing Top Cover.
6. Remove system controller from chassis.
7. Locate system controller card.

8. Push down on ejector levers on each side of system controller until card releases from socket.

9. Grasp top corners of card and pull it out of socket. Place system
controller card on antistatic mat.

10. Using a small flat-head screwdriver, carefully pry battery from system controller.

11. Unpackage replacement battery and Press new battery into system controller with positive side facing upward (away from card).

12. Re-install system controller. Holding bottom edge of system controller, carefully align system controller so that each of its contacts is centered on a socket pin. Ensure that system controller is correctly oriented and ejector levers are open. A notch along the bottom of system controller corresponds to a tab on socket. Push firmly and evenly on both ends of system controller until it is firmly seated in socket. You hear a click when ejector levers lock into place.

13. Use ALOM CMT setdate command to set day and time. Use setdate command before you power on host system.

14. Place top front cover on chassis. Slide front top cover forward until it snaps into place, being careful to avoid catching cover on intrusion switch.

15. Position bezel on front of chassis and snap into place.

16. Open fan door. Tighten captive screw to secure front bezel to chassis.

OBTAIN CUSTOMER ACCEPTANCE (this needs to be performed by the field engineer)
At ALOM prompt:
1. setlocator off
2. Use SC setdate command to set ALOM day and time (before you power on host system).
3. poweron -c

At OK prompt:
Note: The ALOM date and Solaris date are not in sync. The Solaris date may not be correct and should be set separately.
1. Boot -s (to avoid any applications being affected by wrong OS date).
2. Use date command to set the correct OS date (optional: init 0, boot -s to verify)

WHAT ACTION DOES THE CUSTOMER NEED TO TAKE TO RETURN THE SYSTEM TO AN OPERATIONAL STATE:

Customer restarts software applications per applicable administration guides to resume system operation.

Friday, 27 January 2012

How to upload core files to supportfiles.sun.com via ftp

Unsecure File Transfer Options

Oracle recommends that all users employ HTTPS or another secure method to transport files to Oracle. Users that choose to use FTP bear any risk associated with this method of file transport.

Standard method for unsecure upload to supportfiles.sun.com
ftp supportfiles.sun.com
login: anonymous
password: user@machine [ your email address]
cd to the appropriate directory*
binary [ set transfer mode]
put 62001234.tar.Z
quit
*Customers are instructed by the Oracle engineer to choose a destination directory in which to upload their file, based on the customer's location and type of file being uploaded. Choices are:
  • cores
  • iplanetcores
  • explorer
  • explorer-amer
  • explorer-apac
  • explorer-emea

Please Note: Plans are in place to End-Of-Life Supportfiles within the next 12 months. For Oracle Hardware product telemetry and for files greater than 2GB, we recommend Oracle Secure File Transport.

Tuesday, 10 January 2012

How to collect critical troubleshooting information in the SC's log buffer

1) Log into a system which has access to the Main System Controller (SC) and open a terminal window.

2) Open a script session so the following SC command output will be captured.
$ script -a /tmp/scdatafile

3) Connect to the platform shell of the Main SC per your configuration's requirements (telnet, console, ssh, tip, etc):
$ console main-sc
$ telnet main-sc
$ ssh main-sc
$ tip main-sc
    NOTE: Do not reboot the main SC before collecting this data. Doing so may erase critical troubleshooting information in the SC's log buffer.

4) From the platform shell, execute the following commands which will be captured in the script session that you opened previously:
showdate
showsc -v
showescape
showkeyswitch
showcodlicense -v
showcodlicense -rv
showcodusage -v
showplatform -v
showplatform -vda
showplatform -vdb
showplatform -vdc
showplatform -vdd
showboards -ev
showcomponent
showfru -r manr
showchs -b (will fail for fw below 5.20.15)
    And for each suspect or faulty component
showchs -vc /N0/IB6 (for example)
showdate -v
showdate -v -d a
showdate -v -d b
showdate -v -d c
showdate -v -d d
showlogs -v
showlogs -vp (the -vp* commands will fail for systems with older SCs)
showlogs -vda
showlogs -vpda
showlogs -vdb
showlogs -vpdb
showlogs -vdc
showlogs -vpdc
showlogs -vdd
showlogs -vpda
showerrorbuffer
showerrorbuffer -p
showenvironment -ltuv
history
showdate
    NOTE:  You might need to use a "Control right bracket" ("']") to disconnect, depending on how you have connected to the SC.

5) Exit the script session to save the collected data:

    * Hit <control> D and you should get the message "script /tmp/scdatafile closed", "script done" or a similar message.
    * Alternatively, you can also type "exit" at the prompt to close the script session.

6) Upload the data file (scdatafile in this example) utilizing the instructions in Document 1020199.1.

    * It is suggested that the SR Number be apended to the beginning of the file, for example "SR_Number_scdatafile".

Thursday, 1 December 2011

How to Force a Crash Dump When the Solaris Operating System is Hung

In most cases, a system crash dump of a hung system can be forced. However, this is not guaranteed to work for all system hang conditions. To force a dump, you often need to drop down to the boot PROM monitor (OBP) prompt, also known as the "OK prompt", suspending all current program execution.

There are several ways to drop a Sun system to the OK prompt.
1. On older Sun systems with a serial (PS2 type) Sun keyboard and monitor attached, this suspension is performed via a "Stop-A". The upper left key on a Sun keyboard is labeled "Stop". While holding down this key, press the A key.

2. On systems using ASCII terminals for the console, the terminal's predefined break sequence can be used to get to the boot PROM monitor.

3. Newer Sun systems with USB keyboards may require an alternate sequence.

4. Some Sun systems have a system controller/SSP (Enterprise 10000/15000, Sun Fire X800) or ALOM/RSC (Vx80/Vx90 and most new Netra servers) instead of serial port/keyboard access. These can be used to break a hanging system or domain.


Note: There special procedures for Sun SPARC(R) Enterprise Mx000 (OPL) Servers, T1000/T2000 systems, x86 and x64 systems.


The boot PROM monitor will respond with:

Type 'go' to resume
ok

If you don't see this message, you were probably not successful in stopping the system.

Once at the ok prompt, type 'sync' (without the quotes) and press Enter.

The system will immediately panic. Now the hang condition has been converted into a panic, so an image of memory can be collected for later analysis. The system will attempt to reboot after the dump is complete.

The sync command forces the computer to illegally use location, therefore causing a panic: zero. On later revisions of Solaris 8 and above you will see a panic: sync initiated

Not all hang situations can be interrupted. If Stop-A or Break doesn't work, sometimes a series of the same will do the trick. Some hangs are even more stubborn and can only be interrupted by physically disconnecting the console keyboard or terminal from the system for a minute, and then plugging it back in.

If all these attempts fail, you will have to power down the system, thus sadly losing the contents of memory. With luck, a subsequent hang will be interruptable.


NOTE: On the systems with keyswitches, be sure the key is not in the secure position, as this disables the break interrupt in the zs driver.

Sunday, 20 November 2011

zstat-process

Just in-case you encounter a large file /var/adm/exacct/zstat-process.
Here is work-around to reclaim the space.


# df -kh /var
Filesystem             Size   Used  Available Capacity  Mounted on
/dev/md/dsk/d3         4.9G   4.4G       529M    90%    /var

# find /var -xdev -type f -size +100000 -ls -exec du -sk {} \;
  273 2740992 -rw-------   1 root     root     2805391217 Nov 20 06:32 /var/adm/exacct/zstat-process
2740992 /var/adm/exacct/zstat-process

# svcs -a | grep zstat
online         Apr_05   svc:/application/xvm/zstat:default

# svcadm restart svc:/application/xvm/zstat:default

# svcs -a | grep zstat
online          6:38:37 svc:/application/xvm/zstat:default

# df -k /var
Filesystem           1024-blocks        Used   Available Capacity  Mounted on
/dev/md/dsk/d3           5166102     1831942     3282499    36%    /var

# df -kh /var
Filesystem             Size   Used  Available Capacity  Mounted on
/dev/md/dsk/d3         4.9G   1.7G       3.1G    36%    /var

# find /var -xdev -type f -size +100000 -ls -exec du -sk {} \; #

Thursday, 20 October 2011

PICL bug causes Solaris 10 prtdiag to hang

The Solaris PICL framework provides information about the system configuration which it maintains in the PICL tree. I have an experience wherein Solaris 10 prtdiag is hanging. In order to fix this stop and start picld.

# top
load averages: 1582.95, 1462.52, 1345.91 22:57:54
8548 processes:8532 sleeping, 1 running, 1 zombie, 14 on cpu
CPU states: % idle, % user, % kernel, % iowait, % swap
Memory: 8064M real, 2747M free, 4123M swap in use, 9005M swap free
PID USERNAME LWP PRI NICE SIZE RES STATE TIME CPU COMMAND
13622 root 999 59 0 318M 297M sleep 370.3H 82.48% java
26222 root 1 0 0 0K 0K cpu/6 7:52 5.88% ps
25217 root 1 20 0 0K 0K sleep 11:59 5.49% ps
27618 root 1 0 0 0K 0K cpu/5 1:08 5.43% ps
27101 root 1 0 0 0K 0K cpu/4 3:31 5.16% ps

***You can see here that PID 13622 is using alot of CPU.
***And when you check this, it points to prtdiag 

# ps -ef | grep 13622

root 23066 13622 0 Oct 15 ? 0:00 /usr/bin/ctrun -l child -o pgrponly /bin/sh -c /usr/sbin/prtdiag
root 802 13622 0 Oct 15 ? 0:00 /usr/bin/ctrun -l child -o pgrponly /bin/sh -c /usr/sbin/prtdiag
root 28092 13622 0 Oct 15 ? 0:00 /usr/bin/ctrun -l child -o pgrponly /bin/sh -c /usr/sbin/prtdiag

***Restart the PICL
# svcadm restart picl

***Check the load via uptime

# uptime
1:26am up 50 day(s), 11:22, 3 users, load average: 1886.79, 1513.28, 1402.57

***After a couple of minutes check it again
# uptime
1:26am up 50 day(s), 11:23, 3 users, load average: 962.59, 1327.11, 1343.05

***You can observe a dramatic drop on the load
# top
load averages: 3.71, 367.87, 875.60 01:33:23
76 processes: 75 sleeping, 1 on cpu
CPU states: % idle, % user, % kernel, % iowait, % swap
Memory: 8064M real, 5630M free, 1103M swap in use, 12G swap free

PID USERNAME LWP PRI NICE SIZE RES STATE TIME CPU COMMAND
15726 root 1 42 0 108M 33M sleep 21:44 2.89% bptm
11116 root 1 52 0 108M 33M sleep 10:50 0.76% bptm
13622 root 78 59 0 215M 200M sleep 403.3H 0.12% java
15611 root 1 59 0 108M 67M sleep 5:09 0.10% bptm
16953 root 1 0 0 4432K 2160K cpu/11 0:00 0.08% top

Wednesday, 5 October 2011

VERITAS Volume Manager for Solaris

Veritas Volume Manager is a storage management application by symantec ,  which allows you to manage physical disks as logical devices called volumes.

VxVM uses two types of objects to perform the storage management
1. Physical objects - are direct mappings to physical disks
2 . Virtual objects - are volumes, plexes, subdisks and diskgroups.

a. Disk groups are composed of Volumes
b. Volumes are composed of Plexes and Subdisks
c. Plexes are composed of SubDisks
d. Subdisks are actual disk space segments of VxVM disk  ( directly mapped from the physical disks)

1. Physical Disks
Physical disk is a basic storage where ultimate data will be stored. In Solaris physical disk names  uses the  convention like “c#t#d#”  where c# refers to controller/adapter connection, t# refers to the SCSI target Id , and d# refers to disk device Id.  

Physical disks could be coming from different sources within the servers e.g. Internal disks to the server , Disks from the Disk Array  and Disks from the SAN.

Check if the disks are recognized by Solaris

#echo|format
Searching for disks…done

AVAILABLE DISK SELECTIONS:
0. c0t0d0 <SUN2.1G cyl 2733 alt 2 hd 19 sec 80>
/sbus@1f,0/SUNW,fas@e,8800000/sd@0,0
1. c0t1d0 <SUN9.0G cyl 4924 alt 2 hd 27 sec 133>
/sbus@1f,0/SUNW,fas@e,8800000/sd@1,0
 
2. Solaris Native Disk Partitioning

In solaris, physical disks will partitioned into slices numbered as S0,S1,S3,S4,S5,S6,S7 and the slice number S2 normally called as overlap slice and points to the entire disk.  In Solaris we use the format utility used to partition the physical disks into slices.

Once we added new disks to the Server, first we should recognize the disks from the solaris level before proceeding for any other storage management utility.

Steps to add new disk to Solaris:
If the disks that are recently added to the server not visible, you can use below procedure
 
Option 1: Reconfiguration Reboot ( for the server hardware models that doesn’t support hot swapping/dynamic addition of disks )

# touch /reconfigure; init 6

or

#reboot — -r ( only if no applications running on the machine)

Option 2: Recognize  the disks added to external SCSI, without reboot

# devfsadm

# echo | format <== to check the newly added disks

Option 3: Recognize disks that are added to internal scsi, hot swappable, disk connections.

Just run the command “cfgadm -al” and check for any newly added devices in “unconfigured” state, and configure them.

# cfgadm -al
Ap_Id                         Type            Receptacle   Occupant     Condition
c0                                  scsi-bus     connected    configured   unknown
c0::dsk/c0t0d0      disk              connected    configured   unknown
c0::dsk/c0t0d0      disk              connected    configured   unknown
c0::rmt/0                  tape             connected    configured   unknown
c1                                  scsi-bus      connected    configured   unknown
c1::dsk/c1t0d0       unavailable  connected    unconfigured unknown <== disk not configured
c1::dsk/c1t1d0       unavailable  connected    unconfigured unknown < == disk not configured

# cfgadm -c configure c1::dsk/c1t0d0

# cfgadm -c configure c1::dsk/c1t0d0

# cfgadm -al
Ap_Id                         Type            Receptacle   Occupant     Condition
c0                                  scsi-bus     connected    configured   unknown
c0::dsk/c0t0d0      disk              connected    configured   unknown
c0::rmt/0                  tape             connected    configured   unknown
c1                                  scsi-bus      connected    configured   unknown
c1::dsk/c1t0d0       disk              connected    configured unknown  <= Disk configured now
c1::dsk/c1t1d0       disk              connected    configured unknown  <= Disk configured now

# devfsadm

#echo|format <== now you should see all the disks connected to the server


3. Initialize Physical Disks under VxVM control


A formatted physical disk is considered uninitialized until it is initialized for use by VxVM. When a disk is initialized, partitions for the public and private regions are created, VM disk header information is written to the private region and actual data is written to Public region.  During the notmal initialization process any data or partitions that may have existed on the disk are removed.

Note: Encapsulation is another method of placing a disk under VxVM control in which existing data on the disk is preserved

An initialized disk is placed into the VxVM free disk pool. The VxVM free disk pool contains disks that have been initialized but that have not yet been assigned to a disk group. These disks are under Volume Manager control but cannot be used by Volume Manager until they are added to a disk group

Device Naming Schemes
In VxVM, device names can be represented in two ways:

    Using the traditional operating system-dependent format c#t#d#
    Using an operating system-independent format that is based on enclosure names

c#t#d# Naming Scheme
Traditionally, device names in VxVM have been represented in the way that the operating system represents them. For example, Solaris and HP-UX both use the format c#t#d# in device naming, which is derived from the controller, target, and disk number. In VxVM version 3.1.1 and earlier, all disks are named using the c#t#d# format. VxVM parses disk names in this format to retrieve connectivity information for disks.

Enclosure-Based Naming Scheme
With VxVM version 3.2 and later, VxVM provides a new device naming scheme, called enclosure-based naming. With enclosure-based naming, the name of a disk is based on the logical name of the enclosure, or disk array, in which the disk resides.

Steps to Recognize new disks under VxVM control
1. Run the below command to see the available disks under VxVM control

# vxdisk list
in the output you will see below status

    error indicates that the disk has neither been initialized nor encapsulated by VxVM. The disk is uninitialized.
    online indicates that the drive has been initialized or encapsulated.
    online invalid indicated that disk is visible to VxVM but not controlled by VxVM

If disks are visible with “format” command but not visible with  ”vxdisk list” command, run below command to scan the new disks for VxVM

# vxdctl enable

Now you should see new disks with the status of “Online Invalid“

2. Initialize each disk with “vxdisksetup” command

#/etc/vx/bin/vxdisksetup -i <disk_address>

after running this command “vxdisk list” should see the status as “online” for all the newly initialized disks

4. Virtual Objects (DiskGroups / Volumes / Plexs )  in VxVM

Disk GroupsA disk group is a collection of  VxVM disks ( going forward we will call them as VM Disks ) that share a common configuration.  Disk groups allow you to group disks into logical group of Subdisks called plexes which in turn forms the volumes.

Volumes
A volume is a virtual disk device that appears to applications, databases, and file systems like a physical disk device, but does not have the physical limitations of a physical disk device. A volume consists of one or more plexes, each holding a copy of the selected data in the volume.

Plexes:
VxVM uses subdisks to create virtual objects called plexes. A plex consists of one or more subdisks located on one or more physical disks.



Key Points on Transformation of Physical disks into Veritas Volumes
1. Recognize disks under solaris using devfsadm, cfgadm or reconfiguration reboot , and verify using format command
2. Recognize the disks under VxVM using “vxdctl enable“
3. Initialize the disks under VxVM using vxdisksetup
4. Add the disks to Veritas Disk Group using vxdg commands
5. Create Volumes under Disk Group using vxmake or vxassist commands
6. Create filesystem on top of volumes using mkfs or newfs, and you can create either VXFS filesystem or UFS filesystem

Thursday, 29 September 2011

Solaris Troubleshooting – System Panics, Hangs and Crashes

Solaris Troubleshooting – System Panics, Hangs and Crashes

There are a number of differing scenarios under which the Solaris  operating system may panic, hang, or exhibit other symptoms that lead the administrator to have to restart or reboot the system.
As there are many different failure scenarios and many different classes of hardware, the information and procedures for collection of system information vary from system to system.

1.Hang

The first class of failures is the hang. This is when a system appears to become unresponsive. See the documents below that discuss dealing with hung systems.
Be aware that some systems that appear hung are not! Be sure to verify if the server is hung or not. For example: The display may be non-responsive, because the output has been redirected to the console device.

2.Panic or unexpected reboot

Panics can be caused by a variety of issues, including Solaris Bugs, Hardware errors and Third Party Drivers and Applications. It’s important to collect as much detail as is possible when systems panic.
Information that need to be collected to troubleshoot Kernel Panic:
When logging a new case, provide answers to the following questions as an absolute minimum:
  • When did the problem start
  • What changes have been recently made on the system. Important: Anything that has happened since the last reboot is within the scope of this question. Patching, application changes, disk replacements, anything. It’s all important to know when trying to resolve the issues.
  • How often has this failure occurred
  • What may have been going on around the time of the panic and if anything out of the ordinary may have been observed
In the case of a panic or reboot, the messages log and prtdiag output are items that can be quickly sent to Oracle Sun, and that can go a long way towards diagnosis of the cause of the problem, however, an explorer is almost always better.
By far, the simplest way to collect the vast majority of details required to resolve a panic is to collect:
  • Sun Explorer Data Collector output
  • The crash dump
  • Any console messages
a. EXPLORER OUTPUT
If Explorer is run on the system after the incident occurs, it will contain most of what will be required to understand the current configuration of the system, and additional information that may be helpful.
b. CRASH DUMP
If the system generated a system crash dump (check /var/crash/`hostname`), create a compressed tar file containing the unix.* and vmcore.* files and transfer that file to Oracle Sun for analysis. Compressing the tar file (using one of the compress, gzip of bzip2 utilities) reduces the size of the file dramatically and so reduces the time taken to transfer the file to Sun.
c. CONSOLE MESSAGES
Console messages are most important when the server is experiencing hardware issues and the OS is not allowed an opportunity to panic. In cases where multiple unexpected reboots are occurring, and no diagnostic infomation is being provided by the system logs in /var/adm/messages, some form of console logging should be setup as soon as possible to capture the diagnostic information from the console on the next failure.
Recommended NVRAM settings , to Collect Console Messages:
Bring system to OBP level from command line using “shutdown” or “init 0″ commands (either will run all RC shutdown scripts), sync file systems and then drop system to OK prompt. DO NOT use a stop+A key press. The following commands can be executed from the OK prompt or from the command line using the “eeprom <variable=parameter>” command.
at OK prompt # eeprom Description
setenv diag-level max diag-level=max system will run extended POST
printenv boot-device boot-device determine what your boot device is….
setenv diag-device <your boot-device> diag-device=<your bootdev> prevent attempting net boot w diags on
setenv error-reset-recovery sync error-reset-recovery=sync force sync reboot if system drops to OK
setenv diag-switch  true diag-switch =true
reset-all reboot or init 6 system has to reset for changes to take affect

Exceptions
In some cases, it is difficult to collect an explorer.
In the event that explorer cannot be installed or run in a timely manner, the following data is of tremendous value, and should be collected:
a. MESSAGES LOG
Messages logs from /var/adm directory. If there was a panic, the panic message in the file will help determine if we need to analyze a crash dump to diagnose the cause. In many cases a crash dump is not necessary, and waiting for one to be transferred simply increases the resolution time. For example, if there was a hardware reset rather than a panic, the messages log should show that.
b. PRTDIAG OUTPUT
Output of the prtdiag command.
 /usr/platform/`uname -i`/sbin/prtdiag -v
prtdiag gives a summary of a system’s hardware configuration, so that we would know what part to order in the case of a hardware failure. It also gives hardware error messages that can aid in diagnosis.
c. SHOWREV OUTPUT
The output from ‘/usr/bin/showrev -p’ gives a list of the patches installed on the system. This will help eliminate possible casues of the problem and ensure that the correct versions of source code and analysis tools are used during the investigation.

3.Live Dump

On occasion, it is required that a live crashdump be collected. It’s uncommon, as dumping a live system does not capture a completely consistent snapshot of the system. Data is changing while the dump is being written out. Although live dumps are not always consistent, they are still a great source of information for certain types of issues.
collectiongg Live Dump from the Solaris Machine:
1) Before collecting a live kernel dump, a dedicated, NON SWAP, dump device must be configured using dumpadm(1M). The dedicated dump device must not be used in any other way (i.e., filesystem, databases, etc.). The dump device CANNOT be swap or any part of. If the dump device is part of swap, generation of live kernel dump will corrupt the swap area causing the kernel to eventually panic. Also note that any filesystem or data on the dump device disk will be lost.

2)The kernel is running during generation of live kernel dump, and the linked lists that all kernel debuggers use to traverse those linked structures may fail because the list was in flux when saved. Therefore, always run the following ps command to capture the process addresses:
/usr/bin/ps -e -o uid,pid,ppid,pri,nice,addr,vsz,wchan,time,fname
You must use those switches to get addresses with a 64 bit kernel. This will allow you to look at processes since you will have the process address. From there you can generate more complete threadlists, etc. Please see the ps(1) manpage for meaning of those options. For example, the following is what /usr/bin/ps -elf outputs on a 64-bit machine:
> /usr/bin/ps -elf
F S UID PID PPID C PRI NI ADDR SZ WCHAN STIME TTY TIME CMD
19 T root 0 0 0 0 SY 0 Sep 27 0:00 sched
8 S root 1 0 0 41 20 98 Sep 27 0:00 /etc/init -
19 S root 2 0 0 0 SY 0 Sep 27 0:00 pageout
19 S root 3 0 0 0 SY 0 Sep 27 1:05 fsflush
8 S root 267 1 0 41 20 220 Sep 27 0:00 /usr/lib/saf/sac -t 300
8 S root 154 1 0 51 20 325 Sep 27 0:00 /usr/sbin/inetd -s
8 S root 125 1 0 41 20 294 Sep 27 0:00 /usr/sbin/rpcbind
8 S root 48 1 0 47 20 185 Sep 27 0:00 /usr/lib/sysevent/syseventd
8 S root 50 1 0 48 20 160 Sep 27 0:00 /usr/lib/sysevent/syseventconfd
NOTE: The ADDR field is not populated. If you issue the command to capture process addresses as described above, you will see the following output instead:
> /usr/bin/ps -e -o uid,pid,ppid,pri,nice,addr,vsz,wchan,time,fname
UID PID PPID PRI NI ADDR VSZ WCHAN TIME COMMAND
0 0 0 96 SY 10423a60 0 – 0:00 sched
0 1 0 58 20 30000909528 784 30000909848 0:00 init
0 2 0 98 SY 30000908a98 0 10458248 0:00 pageout
0 3 0 60 SY 30000908008 0 104618a0 1:05 fsflush
0 267 1 58 20 30000963530 1760 300009659a8 0:00 sac
0 154 1 48 20 30001708028 2600 30000baaca2 0:00 inetd
0 125 1 58 20 30001709548 2352 30000aa2102 0:00 rpcbind
0 48 1 52 20 300009f0aa8 1480 300009f0dc8 0:00 sysevent
0 50 1 51 20 300009f0018 1280 22d44 0:00 sysevent

Wednesday, 31 August 2011

Sun Hardware Diagnosis at OBP level

It is so common that servers encounter issues related to hardware and those errors cannot be diagnosed by Operating system level utilities.  To perform preliminary diagnosis and to pin point the hardware trouble, system admins have to rely on OBP (Open Boot PROM) diagnosis options.
at OBP level, system admin have three options to investigate the issue
  1. OBP diagnosis commands
  2. OBDiag Outputs
  3. POST Errors

1.Using OBP Commands

Below are the some of the OBP commands that system admins with advanced skill on hardware can use to investigate the trouble.
banner
Displays the power on banner. The banner includes information such as CPU speed, OBP revision, total system memory, ethernet address and hostid.
.enet-addr
Displays the ethernet address
led-off/led-on
Turns the system led off or on.
nvstore
Copies the contents of the temporary buffer to NVRAM and discards the contents of the temporary buffer.
power-off/power-on
Powers the system off or on.
printenv
Displays all parameters, settings, and values
probe-fcal-all
dentifies Fiber Channel Arbitrated Loop (FCAL) devices on a system. 1
probe-sbus
Identifies devices attached to all SBUS slots. Note - This command works only on systems with SBUS slots.
probe-scsi
Identifies devices attached to the onboard SCSI bus. 1
probe-scsi-all
Identifies devices attached to all SCSI busses. 1
set-default parameter
Resets the value of parameter to the default setting.
set-defaults
Resets the value of all parameters to the default settings. Tip - You can also press the Stop and N keys simultaneously during system power-up to reset the values to their defaults.
setenv parameter value
Sets parameter to specified value. Note - Run the reset-all command to save changes in NVRAM.
show-devs
Displays all the devices recognized by the system.
show-disks
Displays the physical device path for disk controllers.
show-displays
Displays the physical device path for frame buffers.
show-nets
Displays the physical device path for network interfaces
show-post-results
If run after Power On Self Test (POST) is completed, this command displays the findings of POST in a readable format.
show-sbus
Displays devices attached to all SBUS slots. Similar to probe-sbus .
show-tapes
Displays the physical device path for tape controllers.
sifting string
Searches for OBP commands or methods that contain string. For example, the sifting probe command displays probe-scsi, probe-scsi-all, probe-sbus, and so on.
.speed
Displays CPU and bus speeds
test device-specifier
Executes the selftest method for device-specifier. For example, the test net command tests the network connection.
test-all
Tests all devices that have a built-in test method.
.version
Displays OBP and POST version information.
watch-clock
Tests a clock function.
watch-net
Monitors the network connection for the primary interface.
watch-net-all
Monitors all the network connections.

 

2.OBDiag


OBDIAG can be used to diagnosis main logic board as well interface boards ( e.g.  PCI /  SCSI / Ethernet / Serial/ Parallel / Keyboard/mouse / NVRAM / Audio /  Video )
To run OBDIAG simply run
OK> obdiag
You can also set up OBDiag to run automatically when the system is powered on using the following methods:
    1. Set the OBP diagnostics variable:              ok setenv diag-switch  true
    2. Press the Stop and D keys simultaneously while you power on the system
Note: On Ultra Enterprise servers, just turn the key switch to the diagnostics position and power on the system, to start obdiag.

 

3.POST

POST is a program that resides in the firmware of each board in a system, and it is used to initialize, configure, and test the system boards. POST output is sent to serial port A  and POST completion status will be indicated by the status LEDs
You can watch POST ouput in real-time by attaching a terminal device to serial port A. If none is available, you can use the OBP command show-post-results to view the results after POST completes.
How To Run POST
  • Attach a terminal device to serial port A.
  • Set the OBP diagnostics variable:ok
ok setenv diag-switch true
  • Set the desired testing level. Two different levels of POST can be run, and you can choose to run all tests or some of the tests. Set the OBP variable diag-level to the desired level of testing (max or min), for example:
ok setenv diag-level max
  • If you wish to boot from disk, set the OBP variable diag-device :
ok setenv diag-device : disk   (  The system default for this variable is net).
  • Set the auto-boot variable
ok setenv auto-boot false
  • Save the changes
ok reset-all
  • Power cycle the system (turn it off, and then back on).
POST runs while the system is powered on, and the output is displayed on the device attached to serial port A. After POST is completed, you can also run the OBP command show-post-results to view the results.
LED STATUS
Power LED ( Left position)
Should always be on. If all three LEDs are off, suspect a power problem. If this LED is in any other state than on and steady, it indicates a problem.
Service LED (Middle Position)
This LED should be off in normal operation. If on, a component is in an error state and you should check check individual board LEDs. A lit service LED does not imply there is an OS-related problem.
Cycling LED ( Right Position)
This LED should be flashing — this is the normal state.

Thursday, 11 August 2011

How to collect a snapshot from an M-Series machine

To collect a snapshot from a Mx000 system you will need the following information:

1.  The name (or IP) of a server on the same subnet as the XSCF
    (or service processor) where you can store the snapshot data.
2.  A user name and password for the server you will be storing the data on.
3.  The full path on the server where you want snapshot to store the data.

The syntax for running the snapshot command on the XSCF is as follows:

snapshot -LF -t username@servername:/full_path_to_data_location -k download

OR you may use the "none" option:

snapshot -LF -t username@servername:/full_path_to_data_location -k none

You will be prompted twice while snapshot is running:

1.  Accept this public key (yes/no)?  Y
2.  Enter ssh password for user '/username/' on host /servername/

*** Once you have gathered the snapshot, please rename the output file to
include your SR number, then upload this file into the /cores directory at
the http://supportuploads.sun.com site. ***

Note:  As an alternative option for collecting a snapshot, you may also direct
the output to a USB memory stick with the following command:

XSCF> snapshot -d usb0

In the event you are unable to collect a complete snapshot the output from the following four XSCF commands can be substituted:

showstatus
showhardconf
fmdump -m
fmdump -V

Sunday, 10 July 2011

Error occurred during initialization of VM

Error occurred during initialization of VM
Could not reserve enough space for object heap

This is an unusual case of Java heap size. At first, we thought that Java is eating up a lot of memory but is not.

According to the application owner the problem happen all the time time, and they had to open up a ticket once in a few month.

Initially I'm not sure if this is actually capacity problem since the initial setting of the java heap size is very small.

My first option is to restart the Weblogic application to free up some memory and see if it’s really eating up a lot of memory.
The application owner agreed to restart the weblogic application.

I also asked if they are you using java console. I had a chance to trace the slope of the increase in memory usage via BMC perform. In the beginning of the slope I found a java console process start-up. I’m almost convinced that it’s the problem maker.

So I asked the application owner to stop that admin console after he restarted the application. But He told me that they are doing everything via command line. So maybe its not the one.

It took about 10 min to come up. But still the application couldn't start up. It’s still saying:

Error occurred during initialization of VM
Could not reserve enough space for object heap

Then I run top command to check what is still eating up the memory.

And I found out that there are sec.pl processes consuming memory.

It is actually the sec.pl which doing a lot of damage with the capacity. As per conversation with the application owner, sec.pl is a monitoring tool they used for the application. As those jobs are very small, and shouldn't take up much we opt to kill these processes.

I found 5 instances of sec.pl eating up more than 500MB each. So upon, killing these processes we got almost 3GB of memory available.

The application was able to run the script after.

Now we are still investigating why these sec.pl processes are eating up a lot of memory.

Friday, 1 July 2011

When to use whole root or sparse root?

It is part of the process in creating a zone to decide whether to use a whole root or the sparse root zone.

Let’s define them first.

A whole root zone is the type of zone which has its own copy of the operating system’s files. On the other hand a sparse root zone shares operating system files with the global zone.

Basically when you need to write on /root, /usr and/or /lib file system you really need to use the whole root zone. This provides the flexibility to customize the instance of your operating system.

In simple terms, when you need to install special software applications into your local zones, you should be creating a whole root zone.

If this is not the case, you may want create a sparse root zone to save up some disk space.




Thursday, 30 June 2011

Where is the mac address located physically?

Here is a quick hint.

For SPARC systems, it is in the NVRAM or the Non Volatile Random Access Memory. The banner command read it from here.
There are times you will find values like ff:ff:ff:ff:ff:ff. This could probably mean that there is a problem reading it from the NVRAM. Perhaps the NVRAM can be defective.
In other SPARC systems, there is a separate chip wherein the mac address is written.

For x86 systems, it is only in the NIC or the Network Interface Card that the mac address is written. No special chip or separate memory holds this information.
Now, what if it's a new machine and you need to know the mac address. Of course you cannot run ifconfig during this occasion. It is easy. Just power on the box and press F12 as it prompts you with bios setup. It will broadcast the mac address. But there is a requirement; the box should be equipped with updated PXE boot environment.


Tuesday, 28 June 2011

How to change timeout value on GRUB

When your x86 Solaris system boots, BIOS loads boot loader (GRUB) from boot device.Then GRUB takes control of the booting.
 
The only issue with this is that the default timeout is only 10 seconds. You may want to decrease or increase this amount.
 
This is quite simple. Just open up the /boot/grub/menu.lst file in your favorite text editor. I’m using vi:

# vi /boot/grub/menu.lst

Now find the section that looks like this:

# menu timeout in second before default OS is booted
# set to -1 to wait for user input
timeout 10

The timeout value is in seconds. You can set it to -1 which stands for user input.
 
Save the file, and when you reboot the change will be set.

How to create ASM links in Solaris hosts

Here is a quick guide in creating ASM links in Solaris boxes.
First thing to do is to check that ASM instance is running. If it is running have the name of disks/LUNs ready.
Do not forget to turn on the MPXIO on Solaris 10 servers by running the following command.

#stmsboot –e

Note that this will reboot the system if the MPXIO is off.

Make sure the disks are not part of any SVM meta devices, VXFS disk group or not in used by any ZPOOL.
Use appropriate commands to verify.

Now, find the physical device path for each disk.
#ls –l /dev/rdsk/|grep <disk name>
For example:
# ls -l /dev/rdsk/ | grep c4t600A0B8000562790000005D04998C446d0
lrwxrwxrwx   1 root     root          64 Feb 16 02:10 c4t600A0B8000562790000005D04998C446d0s0 ->

../../devices/scsi_vhci/ssd@g60060e800542f000000042f000001290:a,raw
lrwxrwxrwx   1 root     root          64 Feb 16 02:10 c4t600A0B8000562790000005D04998C446d0s1 ->

../../devices/scsi_vhci/ssd@g60060e800542f000000042f000001290:b,raw


The idea is to figure out which partition points to what physical device path.

Make sure none of the links under /dev/ora_rdsk are pointing to the disks that you are going to use.
Look for both physical and logical name of the disk.


# ls -l /dev/ora_rdsk/ | grep c4t600A0B8000562790000005D04998C446d0

# ls -l /dev/ora_rdsk/ | grep ssd@g600a0b8000562790000005d04998c446


Now partition the disk in such a way that the slot 0 gets started from the sector 256 to the end. 
Here is the trick, create a temp zpool using the disks.

#zpool create tmp <disk name>

Then destroy the pool.

#zpool destroy tmp

This works better than the format command.

This will create 9 slices for your disk. The first slice (slice 0) is the one that must be used for ASM links.

Create the symbolic links to the physical devise names for the slice 0 of the disks under /dev/ora_rdsk for the requested ASM links.

#ln –s  ../../devices/scsi_vhci/ssd@g600a0b8000562790000005d04998c446:a,raw  /dev/ora_rdsk/VOLUME_NAME

Change the ownership of the physical device to oracle:dba, the default owner is root:sys.

# ls -lhL /dev/rdsk/c4t600A0B8000562790000005D04998C446d0s0
crw-r-----   1 root     sys      118, 64 Feb 16 02:10 /dev/rdsk/c4t600A0B8000562790000005D04998C446d0s0
# chown oracle:dba /dev/rdsk/c4t600A0B8000562790000005D04998C446d0s0 
 
or 
 
#chown  oracle:dba  /devices/scsi_vhci/ssd@g600a0b8000562790000005d04998c446:a,raw

# ls -lhL /dev/rdsk/c4t600A0B8000562790000005D04998C446d0s0
crw-r-----   1 oracle   dba      118, 64 Feb 16 03:00 /dev/rdsk/c4t600A0B8000562790000005D0

You may now ask your DBA to verify.

Monday, 27 June 2011

How to roll back a patch

Let's say you finally realize that the patch you have installed in your box is not doing anything good. Here is a sort roll back instruction. Hopefully you have not attached your secondary mirror yet or else it is too late for this reliever.

Note: The naming conventions defends on you SVM disks naming standards.

Mount the secondary disk

# mount /dev/dsk/cxtxdxsx /mnt


Edit system file on the secondary disk. Comment out the lines that start with rootdev and set md.

# vi /mnt/etc/system

 i.e. set md:mirrored_root_flag=1
       rootdev:/pseudo/md@0:0,0,blk

Backup your vfstab file.
# /mnt/etc/vfstab /mnt/etc/vfstab.md

And edit the file to remove all references to /dev/md devices replacing them with the equivalent non-encapsulated devices.

Then unmount the /mnt and boot from secondary disk.
Then re-create metadevices. Destroy the original mirror and re-create the volumes.
This time the secondary disk as the primary sub-mirror.

# metaclear -r d0
# metaclear -r d6
# metainit d0 -m d20
# metainit d6 -m d26
# metaroot d0


Reboot again from secondary mirror then re-create the metadevice for the primary disk.

# metainit d10 1 1 cxtxdxs0
# metainit d16 1 1 cxtxdxs6

Then re-attach the metadevices to resync.

# metattach d0 d10
# metattach d6 d16

After the resync, your back to the original patch level.



Patching a global zone with more than 3 local zones

Just incase you are patching a global zone with more than 3 local zones. If this is the case, we should detach all of the zones from the container before patching. There is a chance you might encounter some errors while re-attaching resulting to an incomplete state of the zones. One of the cause might be full file system. Make a quick check abd perform cleanup before re-attaching. If it fails, here is a workaround.

globalsys# zoneadm -z zone-sys1 attach -u
Getting the list of files to remove
Removing 5 files
Remove 7 of 7 packages
Installing 594 files
Add 192 of 192 packages
Installation of these packages generated warnings: SUNWzfsr SUNWzfsu SUNWzoner SUNWzoneu
Updating editable files
pkgserv: ERROR: pkglog is not complete
pkgserv: ERROR: Ignoring 5889 bytes from log
pkgserv: ERROR: cannot rewrite the contents file
The file </var/sadm/system/logs/update_log> within the zone contains a log of the zone update.
cat: output error (0/240 characters written)
No space left on device
zoneadm: zone 'zone-sys1': '/etc/release' failed with exit code 2.
could not update zone

globalsys# zoneadm -z zone-sys1 attach -u
zoneadm: zone 'zone-sys1': zone is incomplete; uninstall required.

globalsys# zoneadm list -vic
  ID NAME             STATUS     PATH                           BRAND    IP   
   0 global           running    /                              native   shared
   - zone-sys1          incomplete /zones/zone-sys1                 native   shared
   - zone-sys2          installed  /zones/zone-sys2                 native   shared
   - zone-sys3          configured /zones/zone-sys3                 solaris9 shared
   - zone-sys4          configured /zones/zone-sys4                 native   shared
   - zone-sys5          configured /zones/zone-sys5                 solaris9 shared
   - zone-sys6          configured /zones/zone-sys6                 native   shared
   - zone-sys7          configured /zones/zone-sys7                 solaris9 shared
   - zone-sys8          configured /zones/zone-sys8                 native   shared
globalsys# cd /zones/zone-sys1

globalsys# ls
SUNWdetached.xml  lost+found        root
dev               lu

Move the SUNWdetached.xml file somewhere safe.

globalsys# mv SUNWdetached.xml SUNWdetached.xml_bak.zone-sys1
globalsys# mv SUNWdetached.xml_bak.zone-sys1 /var/tmp


Then make a copy of the /etc/zones/index file.

globalsys# cp /etc/zones/index /etc/zones/index.incomplete


**Now, edit the /etc/zones/index. Change the line containing the "incomplete" keyword and replace it with "installed".

globalsys# vi /etc/zones/index

# Copyright 2004 Sun Microsystems, Inc.  All rights reserved.
# Use is subject to license terms.
#
# ident "@(#)zones-index        1.2     04/04/01 SMI"
#
# DO NOT EDIT: this file is automatically generated by zoneadm(1M)
# and zonecfg(1M).  Any manual changes will be lost.
#
global:configured:/:
zone-sys1:incomplete:/zones/zone-sys1:66e6b6eb-8666-ebe6-ebe8-b6fb6bf66868
zone-sys2:installed:/zones/zone-sys2:f6b66668-6866-eb6e-eeb6-8b66f6bb66fb
zone-sys3:installed:/zones/zone-sys3:ef66e8e6-688b-e6b8-bfbe-866e8f666f6e
zone-sys4:installed:/zones/zone-sys4:66b66b66-68e6-668f-ebe6-eb6bb68efbf6
zone-sys5:installed:/zones/zone-sys5:6e666bf6-8bbb-6688-b6b6-b66eebb666be
zone-sys6:installed:/zones/zone-sys6:6e6b6e6f-6eb8-6b68-b66e-ee686bebbe66
zone-sys7:installed:/zones/zone-sys7:6bf6e6e6-8f6f-66b6-e66e-b6668e66febb
zone-sys8:installed:/zones/zone-sys8:e6666b6b-86ee-e6e6-bbb6-86e6eb66f686

Now, detach the zone to put it in configured state

globalsys# zoneadm -z zone-sys1 detach
globalsys# zoneadm list -vic
  ID NAME             STATUS     PATH                           BRAND    IP   
   0 global           running    /                              native   shared
   - zone-sys2          installed  /zones/zone-sys2                 native   shared
   - zone-sys1          configured /zones/zone-sys1                 native   shared
   - zone-sys3          installed  /zones/zone-sys3                 solaris9 shared
   - zone-sys4          installed  /zones/zone-sys4                 native   shared
   - zone-sys5          installed  /zones/zone-sys5                 solaris9 shared
   - zone-sys6          installed  /zones/zone-sys6                 native   shared
   - zone-sys7          installed  /zones/zone-sys7                 solaris9 shared
   - zone-sys8          installed  /zones/zone-sys8                 native   shared

That's it. Try to reattach the zone again. It should be good by now.

globalsys# zoneadm -z zone-sys1 attach -u
Getting the list of files to remove
Removing 5 files
Remove 8 of 8 packages
Installing 6 files
Add 9 of 9 packages
Updating editable files
The file </var/sadm/system/logs/update_log> within the zone contains a log of the zone update.
globalsys# zoneadm list -vic
  ID NAME             STATUS     PATH                           BRAND    IP   
   0 global           running    /                              native   shared
   - zone-sys2          installed  /zones/zone-sys2                 native   shared
   - zone-sys1          installed  /zones/zone-sys1                 native   shared
   - zone-sys3          installed  /zones/zone-sys3                 solaris9 shared
   - zone-sys4          installed  /zones/zone-sys4                 native   shared
   - zone-sys5          installed  /zones/zone-sys5                 solaris9 shared
   - zone-sys6          installed  /zones/zone-sys6                 native   shared
   - zone-sys7          installed  /zones/zone-sys7                 solaris9 shared
   - zone-sys8          installed  /zones/zone-sys8                 native   shared