Showing posts with label SAN. Show all posts
Showing posts with label SAN. Show all posts

Monday, April 13, 2026

NDFC upgrade / deploy part 3 - deploy new NDFC one node setup in ESX

 As mentioned in previous post, we choose 1 node setup with App OVA with 16vCPU and 64GB memory.  It will be run in ESX cluster.  

Pls review the link below 

Deploying Nexus Dashboard Using VMware vCenter

Cisco Nexus Dashboard Fabric Controller 12 

For SAN, you still need the following ready before deployment.
1) 2 NICs: one for management IP and one for data IP and both can be in the same subnet.  
2) DNS and NTP are required
3) Cluster name and VM name should be different but it is a one node setup.  So, it does not really matter.
4) Pool IP: 2 are required in our setup (2 IPs for data network)  Ask for an additional IP if you plan to use  SAN Insight).  All pool IP should be whitelisted for SMTP traffic in additional to Management and Data IP.  
5) Internal subnets settings (these subnets are used internally and will not leave cluster.  Just in case, confirm the subnets are not in use and subnet mask has to be 255.255.0.0)
App subnet and Service subnet
6) May want extra space for HDD depends on the number of switches to manage.

The steps are very straight forward.  However, depending on the environment, do not specify VLAN in step 18 for data network.  Otherwise, the data network will not be able to communicate with the switches even they are in the same subnet.  In our environment, VLAN info is already present at the switch and ESX level.  

For initial setup, you can follow the document below 

Feature Management will be set to SAN Controller with Performance Monitoring and SAN Web Device Manager for my environment.  









Server settings are self explanatory.  For SMTP, make sure all pool IPs are added to SMTP allowed list.  

If you want to confirm SMTP alert, go to Event -> Event setup -> Forwarding -> Add Rule (same as DCNM).  Multiple recipients allowed and event email will be sent thru pool IP.  

Below link provided detail steps on how to add event email.

How to Configure Event Forwarding using Cisco Nexus Dashboard Fabric Controller - Cisco

I include a blog post below for bugs we faced for NDFC appliance.  See 

NDFC bugs since deployed

Monday, March 2, 2026

NDFC upgrade / deploy part 2 - sizing and compatibility

At the time we were planning, we decided to use ND 3.2.1i (which comes with NDFC 12.2.2).  That would work with all our existing linecard, switches and firmware.  I don't need to worry about existing firmware 8.4(2e) was not supported.  Please use links below for compatibility check.  

For software / hardware compatibility, pls check NDFC Software and Hardware Compatibility Matrix.

To confirm the NDFC and ND compatibility, pls check Cisco Nexus Dashboard and Services Compatibility Matrix    

Next is sizing.  For sizing, pls check Cisco Nexus Dashboard Capacity Planning.  We decide to use 1 App node config since it is a small environment.  (Note: if one node config is decided, it will not support adding any more node in the future).  App OVA requires 16 vCPU and 64GB memory.  

Other requirement: pls go through documents below

Cisco Nexus Dashboard and Services Deployment and Upgrade Guide, Release 3.2.x - Prerequisites: Nexus Dashboard [Cisco Nexus Dashboard] - Cisco


Device Manager error in Cisco

Recently, when we try to open Device Manager to make changes on the MDS 9710 switch, the error below pops up.  It only happens to one of the sites.  The other site works fine.  We try to open device manager thru NDFC and get the same error.  





We did see a different error in the past and needed to switchover the controller in the past.  The error was "Busy network, no route, or snmpd is unresponsive."

See EMC kb 000218149.  However, it still fails to open.  

We suspect something with network but still open a ticket with Cisco.  Support runs some trace and does not see issue on the switch side and there is connectivity between switch and the client running Device Manager.  Eventually, we workaround the issue by confirming TCP is set to true for DeviceManager.bat.

set JVMARGS=%JVMARGS% -Dsnmp.preferTCP=true

We still have trouble after that.  So, we login to the switch in the other site that we don't have issue with.  Then click Device > Preference.  






Select TCP for Use SNMP on next launch. Click Apply then Ok.  

Try to use Device Manager on the switch that we have issue.  Now it works fine.  Only problem with the workaround is to have a switch that can connect to Device Manager without issue.  I have not asked support if there is any other way to force it to use TCP instead of UDP.  


NDFC bugs since deployed

NDFC was deployed in last Oct.  Couple of bugs were discovered.  

1) About 4 - 6 weeks, we will see an alert 

Elasticsearch error - 'could not fetch component status'





The bug ID is CSCwm51621.

If you follow the bug ID above, there is a solution from the forum.  For me, when that error comes up, I just reboot the appliance with "acs reboot".  The error will go away, and it won't come up until another 4 - 6 weeks.  A normal reboot is sufficient and DO NOT add any other option after "acs reboot".  

2) /logs/k8/pods 90% usage alert.  About 4 months after deployed, an alert /log/k8/pods 90% usage showed up in the Admin Console of the Nexue Dashboard.  







Support was contacted to clear the old logs.  Currently, there is no fix, and webex is required for support to clear the old logs manually.  In the future release, the log retention for the folder will be changed based on the chat with support.  

Tuesday, December 30, 2025

NDFC upgrade / deploy part 1 - License planning

With DCNM reaching EOL in Apr 2026, the only option left is to upgrade to NDFC.  Some site may decide to run with DCNM until the whole infrastructure is migrated to cloud.  Because we are still refreshing old MDS 9700 switches, migrating to NDFC is a must.  The main issue with firmware upgrade for existing switch is smart license.  We manage no more than 10 MDS FC switches at a time.  So, setup is kind of simple.  Still, it is more complicated compared to DCNM.

It looks like 9.2(1a) is the last version of firmware in MDS 9000 supporting legacy license / license file.  If you upgrade existing switch running older version of firmware to 9.2.2, the switch license will be changed to smart licensing auto.  Check with support.  The answer I got is those existing switch upgraded to newer firmware will be in Donor mode in the worst case.  There should be no impact to functionality according to support.  Because some of the MDS switches licenses were purchased through 3rd party vendor, and we change to another 3rd party vendor for support later, I don't know what will happen if there is licensing issue after firmware upgrade on existing switch.  So, we decide to keep existing older MDS 9700 switches with firmware 8.4(2e) until hardware refresh completes in the next 2-3 yrs.  There are a few bugs impacting firmware 8.4.  Confirm you have all the workarounds before deciding to stay in 8.4 firmware.  If you decide to upgrade existing MDS 9700 to use smart license, check with support first.   

For all the new MDS 9700, they are all shipped with firmware 9.4.x.  So, smart license is enabled auto.  Check with support and there is no license file any more on these new MDS 9700 switches.  License is installed at the factory.  

Because it is a close environment, we setup our NDFC smart licensing to offline mixed mode.  Mixed mode will support existing license file in older MDS 9700 with older 8.4 firmware.  We contact support to transfer existing legacy DCNM server license to smart license before upgrade.  At most 365 days / switch addition or decommission, I will export the license info back to Cisco software support website and then import it back to the NDFC server licensing section.  

Below are some of the Cisco Smart licensing links / doc for your information.  

Brownfield_Conversion_QRG

Cisco MDS 9000 Series Licensing Guide, Release 9.x - Smart Licensing Using Policy [Cisco MDS 9000 NX-OS and SAN-OS Software] - Cisco

Cisco MDS Smart Licensing Using Policy Data Sheet


Saturday, April 19, 2025

DCNM login error "RequestSendFailed: EJBCLIENT000409" for the SAN Client

We have plan to build a new NDFC to manage our MDS switches.  It is more complex and will take some time to deploy.  Few challenges are resource requirement, license transfer and other new requirement.  In the meantime, we continue to use existing DCNM 11.5(4).  

Last week, suddenly, there is a login error with the java SAN client "RequestSendFailed: EJBCLIENT000409".  We have restarted the services and reboot the DCNM server.  However, it does not help.  





Do some research and most likely it is related to certificate.  We don't have our own certificate.  So, most likely, the default expired.  It is confirmed by checking the certificate expiration date from the web browser.  So, open a ticket with support and ask them to renew the certificate for 2 more years.  Then, I can login without issue.  

Monday, June 5, 2023

No Port Down email alert from DCNM

We used to have SNMP to monitor all FC port down event.  However, it causes a lot of un-necessary calls from NOC after hours.  Because we have most of the prod servers setup with 4 HBAs, we decide to use SMTP alert to email to storage team members only.  However, it never seems to work with DCNM.  We did see Port up alert but not for Port down regardless of what settings we choose.  

After opening a ticket with support, we were told the setting was disable by default.  If you want to receive email alerts for port down event, go to Server Properties of DCNM.  Look for event.linkDown.log and set it to True.  By default, it is set to False.  Then restart the DCNM services and port down alerts will be sent to your SMTP server depending on your setup.  



Cisco bug CSCvz61883

Check the bug info at Cisco site for CSCvz61883 / EMC article 000197332 and it applies to the MDS 9700 32G FC modules in our environment.  The DS-X9648-1536K9 linecard in MDS 9700 can be affected at almost exactly 468 days of uptime.  

Our MDS 9710 with 32G linecard are running 8.1(1a) with supervisor 3 and 8.4(1a) with supervisor 4.  Runs the command below against the 32G linecard in MDS9710 to see module uptime.  In our environment, they reach 426 and 417 days respectively.    

slot x show system uptime   

8.4(2d) does provide the fix to the issue but it is quite new.  Open a ticket with support and the recommendation is to upgrade to 8.4(2c).  This will reset the module uptime back to zero and buy another 400+ days.  With 8.4(2), there is a workaround to reset the uptime for the device non-disruptively.  You can consider another firmware update after 400 days on the FC switch.  

Finished the upgrade to 8.4(2c) last week.  Upgrade to 8.4(2c) on MDS 9710 with supervisor 3 is much faster than MDS 9710 with supervisor 4.  

=================================================================

Updated in late May 2023

Another reason we never considered 8.4(2d) was due to CSCwb29379 (see release notes 8.4(2e) below) 

Cisco MDS 9000 Series Release Notes, Release 8.4(2e)

8.4(2e) was available for quite some time now.  So, we upgraded to 8.4(2e) couple weeks ago to fix this problem permanently.  

Sunday, February 12, 2023

Verifying that Every EMC Disk Arrays are Properly Attached to an EMC SMI-S Provider

Just update ViPR SRM to point to the new Windows 2016 with SMI-S provider.  However, SMI-S complains it cannot retrieve the info of the array.  Confirm ECOM is running and symcfg list shows the VMAX array.

Restart ECOM service and ViPR SRM can retrieve the stats from the array.   However, I try to find out how to confirm EMC SMI-S does detect the array. 

Follow the kb.
https://www.sentrysoftware.com/kb/KB1132.html

Below is the kb copied from the link. 

Objective

Monitoring EMC disk arrays requires to configure the EMC SMI-S Provider to make sure that the arrays are being discovered. EMC disk arrays are automatically discovered when they are locally attached to the EMC SMI-S Provider. This article explains how to make sure that the EMC SMI-S Provider discovers all disk arrays by using the TestSmiProvider tool.

Procedure

  1. Logon to the server where the EMC SMI-S provider is installed
  2. Run the TestSmiProvider.exe which can generally be found under “C:\ProgramFiles\EMC\ECIM\ECOM\bin”
    C:\Program Files\EMC\ECIM\ECOM\bin>TestSmiProvider.exe
    Connection Type (ssl,no_ssl) [no_ssl]:
    Host [localhost]:
    Port [5988]:
    Username [admin]:
    Password [#1Password]:
    Log output to console [y|n (default y)]:y
    Log output to file [y|n (default y)]:
    Logfile path [Testsmiprovider.log]:
    Connecting to localhost:5988
    Using user account 'admin' with password '#1Password'
    ########################################################################
    ##                                                                    ##
    ##                  EMC SMI Provider Tester                           ##
    ##   This program is intended for use by EMC Support personnel only.  ##
    ##   At any time and without warning this program may be revised      ##
    ##   without regard to backwards compatibility or be                  ##
    ##   removed entirely from the kit.                                   ##
    ########################################################################
      slp    - slp urls                     slpv    - slp attributes
      cn     - Connect                      dc      - Disconnect
      disco  - EMC Discover                 rc      - RepeatCount
      addsys - EMC AddSystem                remsys  - EMC RemoveSystem
      refsys - EMC RefreshSystem
      ec     - EnumerateClasses             ecn     - EnumerateClassNames
      ei     - EnumerateInstances           ein     - EnumerateInstanceNames
      ens    - EnumerateNamespaces          mine    - Mine classes
      a      - Associators                  an      - AssociatorNames
      r      - References                   rn      - ReferenceNames
      gi     - GetInstance                  gc      - GetClass
      ci     - CreateInstance               di      - DeleteInstance
      mi     - ModifyInstance               eq      - ExecQuery
      gp     - GetProperty                  sp      - SetProperty
      tms    - TotalManagedSpace            tp      - Test pools
      ecap   - Extent Capacity              pd      - Profile Discovery
      im     - InvokeMethod                 active  - ActiveControls
      ind    - Indications menu             tv      - Test views
      st     - Set timeout value            lc      - Log control
      sl     - Start listener               dv      - Display version info
      ns     - NameSpace                    vtl     - VTL menu

      q      - Quit                         h       - Help
    ########################################################################
    Namespace: root/emc
    repeat count: 1
  3. Run the command eq at the prompt
    (localhost:5988) ? eq
    Query Language[DMTF:CQL]:
  4. Run the below query
    Query []: SELECT EMC_ArrayChassis.SerialNumber FROM EMC_ArrayChassis

A working provider will return the following for each disk system attached to this provider:

++++ Testing ExecQuery:  ++++
Instance 0:
ObjectPath : //10.0.10.54/root/emc:Clar_ArrayChassis.CreationClassName="Clar_ArrayChassis",Tag="CLARiiON+CKM00083900053"


Clar_ArrayChassis


CLARiiON+CKM00083900053


CKM00083900053

Number of instance qualifiers: 0
Number of instance properties: 3
Property: CreationClassName Number of qualifiers: 0
Property: Tag Number of qualifiers: 0
Property: SerialNumber Number of qualifiers: 0
ExceQuery 1 instances; repeat count 1;return data in 0.000000 seconds
Retrieve and Display data - 1 Iteration(s) In 0.062400 Seconds

A failing provider will return the following:

++++ Testing ExecQuery:  ++++
Error: Connection closed by CIM Server.
Retrieve and Display data - 1 Iteration(s) In 0.053000 Seconds
Please press enter key to continue...

Friday, August 26, 2022

How to find HBA WWPN in Windows 2008, 2012 and 2016

Checking a lot of site and see others suggest to use fcinfo.  For Windows 2008, you can use Storage Explorer.  It will not only show you the WWPN on the servers, but also the zoneset if it is connected to Cisco FC switch. 


For Windows 2012 and 2016, I see other option like powershell.  However, there is a command built-in which is very convenient.  Just run get-initiatorport and it will return the WWPN and WWNN. 


=========================================================================
Update doc 2022 Aug to include Windows 2016

Thursday, April 21, 2022

Update firmware for MDS 9700 switches

 You can find the steps at Cisco site to perform non-disruptive upgrade on MDS 9000 switches.  I add some additional steps.  

1) Read release notes and discuss with support to determine which version to upgrade.  Then, download the correct firmware file.  Supervisor 3 and supervisor 4 in MDS 9710 firmware files are different.  

2) Confirm running config is applied to startup config

copy running-config startup-config

3) Copy running config to bootflash

copy running-config bootflash:$(SWITCHNAME)-$(TIMESTAMP).cfg

4) Backup config, existing kickstart and firmware to tftp.  Other than startup config, I make sure I have a copy of the current version of kickstart and firmware in the TFTP server.

copy bootflash: tftp:

5) Save a copy of show tech-support detail.  In case something goes wrong, support can compare the switch condition before and after upgrade.  

6) confirm bootflash free space before uploading new firmware.  I will delete firmware and kickstart file older than the current version.  
dir bootflash:
dir bootflash://sup-standby/ 

7) upload image and kickstart files using tftp
copy tftp: bootflash:

8) Verify the MD5 checksum

show file bootflash:filename md5sum

9) Run the command below to see the impact of the upgrade.  
show install all impact kickstart bootflash:kickstart.bin system bootflash:image.bin (check if upgrade is non-disruptive)

Run commands below for basic health check
show system redundancy status 


























show module





















10) Confirm no TFTP / SFTP session (CSCvu52058)
In order to see if there are open file transfer sessions, you can run

show users | inc ssh | wc l

and
show processes | inc dcos_ssh | wc l


to compare the number of ssh users logged in and the number of sshd processes running.

If the values are different, you may have an open file transfer connection.

11) Run command below to see if the existing config incompatible with image.  
show incompatibility system bootflash:image.bin

12) Run command below to start upgrade
install all kickstart bootflash:kickstart.bin system bootflash:image.bin 
Once again, you see the impact of the upgrade.  













It lists the current and target version of modules















Then, type y and enter to proceed the upgrade.

Wait for the switchover message


















Login again thru putty.  Now, the active director is module 6.

To show installation process.  Run the command below.
show install all status












































Once it shows "Install has been successful", run the command below to confirm new version of firmware is applied to the switch as expected.  

show version
show module


Wednesday, November 10, 2021

PowerMax: embedded Unisphere with Solution Enabler

Just finish setup of PowerMax and decide to use the embedded Unisphere instead of installing one.  Besides, you can use it as SE server too.  You will need to complete the steps below.  Same info can be found in "Dell EMC PowerMax and VMAX All Flash: Embedded Management"  or 

1) Add the Solution Enabler client IP and user to the Nethost settings (don't remove the default ones)


2) Have support to check the certificate matching the FQDN of the management port if there is issue
3) Make sure SYMAPI server daemon is running.  If it is not running, engage support to start the daemon.  Unfortunately, there are restriction on what you can do on the embedded Unisphere.  
4) Make sure Solutions Enabler Base Configuration > Use Access ID, set this value to ANY.  Again, you need support to dial in to make the change.


5) Edit netcfg in SYMAPI directory in the SE Client.  Add the entry of both management IP and primary IP first.  In the case below, if you manage two PowerMax, you can add the 2nd site to netcfg file.  
SITEAEMGMT Ordered TCPIP FQDN 10.1.1.2 2707 SECURE
SITEAEMGMT Ordered TCPIP FQDN 10.1.1.3 2707 SECURE
SITEBEMGMT Ordered TCPIP FQDN 10.2.1.2 2707 SECURE
SITEBEMGMT Ordered TCPIP FQDN 10.2.1.3 2707 SECURE

If you need to manage only one PowerMax, you can setup the env var.  If you need to manage multiple PowerMax, create a batch file for site A and another one for site B.

set SYMCLI_CONNECT=SITEAEMGMT
set SYMCLI_CONNECT_TYPE=REMOTE
symcli -def
symcfg list


Thursday, May 13, 2021

esrshttpdlistener keeps stopping

 After rebooting eSRS GW, it shows it is online but it shows VE is not connecting to EMC.  Checking services and see esrshttpdlistener is not started.  Try to restart service and it won't stay up.  Strange part is all network test are fine.  

Most of the eSRS is firewall related.  If it passes the connectivity test, check the hosts file as some users point out in the forum  https://www.dell.com/community/Secure-Remote-Services-ESRS/ESRS-3-06-esrshttpdlistener-keeps-stopping/td-p/7039948/page/2

Compared the hosts file with the other gateway, I found out the one with issue has an entry 

127.0.0.1  GW_hostname_longname  GW_hostname_shortname localhost

So, I modify the entry to

127.0.0.1 localhost

Reboot and everything comes up fine.  

Sunday, December 27, 2020

Unisphere 9.x with VMAX 40k

Recently, I update Unisphere and Solutions Enabler to 9.0 for the VMAX 40k.  Unisphere 8.4 will not be supported after May 2020.  Few things are not supported for VMAX 40k in Unisphere 9.0.  Belows are the things I don't like about new Unisphere 9.


  1. I cannot export list of luns in storage group.  I guess EMC wants customer to use ViPR SRM to run report (not sure if this is only applied to VMAX 40k)
  2. No more provision template for VMAX 40k and that is not convenient.  
  3. Cannot change FAST priority in the GUI any more for SG of VMAX40k
  4. Heatmap is not shown in a single page.  
Heat Map in 8.4

Heat Map in 9.0


      5.   In Unisphere 8.4, if you want to delete meta members after dissolve a meta lun, you can select the option to delete meta members.  However, it will only delete meta head in Unisphere 9.  It will be fixed in Unisphere 9.1 according to support.  Because there is no more meta from VMAX, I suspect the developer forget about the VMAX 40k.    


6. There is a limitation with Unisphere 9.0.  Only the first 1024 luns will be shown.  You can use the filter to display the luns after the first 1024.  The filter for lun label does not work too well in 8.x.  However, once it is fixed in version 9, now, only the first 1024 luns will be shown.  See EMC kb Unisphere for PowerMax: Maximum viewable amount is 1024 volumes(000529643)

There was a bug in earlier version of Unisphere 8.x.  When you try to delete a meta lun under Virtual lun, it will crash the Unisphere GUI for VMAX 40k.  It is fixed in Unisphere 8.4.  Glad it is passed to Unisphere 9.0.  Also, if you try to create metalun from Unisphere, striped is populated auto and cannot be changed.  I have not tried to create concatenated lun with CLI in SE 9.0.

Just a reminder don't forget to change the smc password after installing Unisphere.



For VMAX 40k, you cannot change the Unisphere smc password running in SP.  So, make sure to restrict IP access to SP from the network. 

========================================================================
Just upgrade to Unisphere 9.1 (12/2019) 
It does fix the issue mentioned before for dissolve and delete meta luns.  However, it generates 3 new problems.

a) AD login does not work after upgrade.  (using local account at the moment)
b) Cannot customize performance alert for each thin pool / disk group.  (no workaround)
c) Dash Board reduces the overall health by 20 points because SSD pool is 95% full (recommendation  by vendor is to reserve only 1% of space given no lun is bounded / pinned to SSD and only thin provision is used).  This is not an issue.  Just ignore it. 

Hopefully, next patch fixes it. 

Also database in 9.1 is different.  Upgrade fails at database upgrade.  Since I ran all the monthly reports, I actually need to uninstall 9.1 and then install again in another folder.  So, it is a fresh installation for me.  
========================================================================
Update to Unisphere 9.1.0.20 (11/2020)  

AD login issue is still not resolved.  Still using local account.  

Alert customization for each Thin Pool / Disk Group seems fixed if I set it up in IE (won't work with Chrome).   

I doubt the last one mentioned before will be fixed.  Dash Board still reduce overall health by 20 points for SSD usage over 95%.  I will probably stay in this version until migration to PowerMax.  

Friday, August 14, 2020

AppSync login page won't load with 404 error

 After Windows patching and AppSync server reboot, the management page of AppSync won't load.  It only show the 404 error Not found.  Restart server makes no difference.  Check from support page and find the kb 501522  https://support.emc.com/kb/501522

I do see the three files with undeployed below.  So follow the kb and the issue is fixed.  

apollo.ear.undeployed

archway-ear.ear.undeployed

remotex.rar.undeployed

Below is kb501522 from EMC support site.  

Cause

This may be seen if the Appsync Server services are Started and Stopped in quick succession. There are a number of files, such as E:\EMC\AppSync\jboss\.\standalone\..\applications\apollo.ear which are deployed by the Appsync Server during startup. If the service is stopped before these files are fully deployed the Appsync Server may begin to throw this error.

When we check the location C:\EMC\AppSync\jboss\applications we should see the following 6 files in a healthy system: 

  • Apollo.ear
  • apollo.ear.deployed
  • archway-ear.ear
  • archway-ear.ear.deployed
  • remotex.rar
  • remotex.rar.deployed


When we check on a system showing the "Error 404 not found" error we should find some of these files marked as undeployed, eg "apollo.ear.undeployed ".

Workaround

In order to resolve this issue we have to get the Appsync Server service to deploy the .EAR files. Follow the procedure outlined below to accomplish this. 
When we find undeployed files in the C:\EMC\AppSync\jboss\applications folder perform below mentioned steps to resolve:

  1. Stop all the services.
  2. Take backup of the C:\EMC\AppSync\jboss\applications.
  3. Rename Apollo.ear.undeployed to Apollo.ear.dodeploy.
  4. Perform the same steps as step 3 in case any other EAR files are also shown as undeployed.
  5. Start the services.
  6. Check the status for .EAR files again at the location C:\EMC\AppSync\jboss\applications.
i. Apollo.ear
ii. apollo.ear.deployed
iii. archway-ear.ear
iv.archway-ear.ear.deployed
v.remotex.rar
vi.remotex.rar.deployed
  1. Try to log in into appsync GUI after 5 minutes and you should be able to log in.

Tuesday, January 21, 2020

Update Cisco DCNM to 11.3(1) due to security Vulnerabilities

Because of Cisco DCNM security issues (see below), just update DCNM to 11.3(1).  Update should be done ASAP.

I am not using the appliance and it is running on Windows with version 11.1(1).  It is running with Oracle Express 11g and managing only MDS FC switches.  We don't use advance features like SAN Insight and it is a standalone Windows server.  Upgrade step is pretty straight forward.

https://www.cbronline.com/data-centre/cisco-data-center-network-manager/


Thursday, October 31, 2019

Equallogic dirty cache

I used to manage a large number of Equallogic 6510.  They have bought them for production used because of insufficient budget for enterprise storage.  It is just an entry level array with very limited redundancy.  Controllers are active / standby and it takes 21 s to failover when it is completely idle.  So, you can imagine how long it will take if there is heavy IOPs.

It is a pain to update firmware since it takes some time for controller to failover.  Linux and Unix will not like that.  If you used it for production, it will be really hard to get downtime.

One time, there was a bug and both controller panic.  After it starts back up, the management interface is not reachable.  Check serial connection and see the following.  Even though the controllers panic, it still displayed a msg "This is a POWER FAILURE RECOVERY".  Also, it showed RAID LUN not recoverable.  Keep looking down, sounds like we actually have a dirty cache stuck in the memory.  This normally happens with power failure.  


When I try to login from serial port, it shows the array is not even configured.  Obviously answer No when asked to config the array.  


Contact support from that point and it is indeed a dirty cache problem.  Talk to support and confirm if the dirty cache stuck, user will lose management interface access.  Also, only later model of EQL support port failover to standby controller.  If both interfaces from the active controllers die, the interface will not failover to standby controller interface.  You will need to do a manual failover.  That's why I don't suggest them for production. 

Clear the cache and reboot the array.  Everything is normal from that point.  Only thing you lost is the data in the stuck cache.  Luckily, there is no database running in those arrays.  The uncorrectable sectors are empty space. 

Earlier version of firmware especially version 5 and 6 are problematic.  Lots of problem.  After the latest patch of version 7 installed, we see stability from that point.  However, new enterprise arrays were installed, and these units were used for backup / archiving.  Now, they were all retired.  





Sunday, September 8, 2019

AppSync with VMAX 40k

AppSync 3.5 was installed couple years ago to protect our SQL application.  Basically, AppSync agent was installed in SQL cluster.  The SQL clusters were running as VMs and the databases were residing in RDM.

Kept getting VSS error that it takes more than 10s to create VSS.  That was what MS supported for VSS.  If it took more than 10s, the VSS creation step would fail.

Installed latest Solution Enabler 8.4.x.x available at the time and no change.  Since the array was VMAX 40k, DNS was not a factor.  Eventually, version Version 3.5.0.1_URM00111091_PRELIM_R2  fix the bug VSS 10s delay bug with VMAX 40k.

Couple things I don't like about the AppSync.
1)  Mounting and Dismounting SQL VMFS takes long time and does not work well.
2)  Need to reserve an extra copy of luns in the pool   
3)  If info is not sync, AppSync does not know what to do.  So, manually dismount copy from mount host will cause problem because AppSync does not know the luns are dismount.  Eventually, contacting support to clean up is required.

------------------------------------------------------------------------------------------------------------

Now, 3 yrs later, AppSync 3.9 was setup for proof of concept few months ago.  This time, we have SQL running in VMFS.  We encounter same issue as before.  Mounting copy to mount host is very slow and timeout.  Eventually, we keep our design for SQL as before in RDM.  Things are working smoothly with RDM as expected.   

We are still using the AppSync 3.5 now until we migrate to AppSync 3.9 next yr.  Recently, SQL team change the DB structures and instead of a few larger database, we have a lot of smaller database to be protected.  The AppSync host plugin service has memory leak problem.  So, we have to setup a process to bounce the service once every 3 days.  Hopefully, this won't happen in version 3.9.

Thursday, August 29, 2019

DCNM version 11.1(1) and bug ID CSCvf99665

Recently build a new DCNM 11.1(1) box in Windows 2016 to replace the existing 7.2(3) because the old one is running Windows 2008 R2.  Major difference is HTML5 and the webclient is a lot faster and most of the work even port channel can be completed in the GUI (I have not tried that yet).  If you don't like to use the new webclient to complete your zoning, you can still use the old FM. 

Besides, I use the Oracle Express for the db of DCNM since very 7.2.  The performance is better than the POSTGRESQL.  

You can follow the Oracle link below to have some basic knowledge of Oracle Express.

If you use SolarWind as your TFTP server, make sure .NetFramework 3.5 is required.  See link below on how to enable it in Windows 2016 server.

After new DCNM server is in production, I plan for the firmware update on all the FC switches in the fabric.  However, I find out both of the MDS9710 are affected by Cisco bug ID CSCvf99665

It show an invalid IPV6 IP address in the mgmt/0 interface and has a zero length subnet mask.

Example:
::148.237.143.255/0

Suggestion from support
(1)Open a case and get Cisco TAC to send you a DPLUG file that will be downloaded and run on the switch.  We would run the DPLUG, do a 'copy r s' to save the configuration, then do a 'system switchover' and run 'copy r s' again.

(2)Upgrade to 8.1(1a).  After the upgrade has completed, do a 'copy r s',  do a 'system switchover' and after both supervisors are back up, then do another 'copy r s'.

On the safe side, we choose the 1st option and have support apply the DPLUG file.  Everything is fine.  Then, we update the firmware to 8.1(1a).  

Sunday, July 24, 2016

ESX host does not discover new path after migrating to new fabric

Once, I work on fabric migration.  They want to migrate to new equipment and choose not to link the existing and new switch together (no ISL) to have a clean config in new fabric.  The servers connect to different fabric A and B.  On each migration, we need to re-connect storage controller on one fabric and the servers that are connecting the controller to new fabric at the same time.  Only one fabric and one controller will be worked on each time.  Since they are all ESX hosts and no AIX, it is much easier. We complete the job in the weekend during low I/O period.  What happens is after connecting to the new switch, it does not see the new path even VI team rescan.  I did confirm zoning had been completed correctly on the new fabric.  They suggest me to follow VM article Configuring fibre switch so that ESX Server doesn't require a reboot after a zone set change


However, I know RSCN is enabled in the Cisco switch.  Eventually, VI agree to vMotion the VMs in the host, and reboot the ESX host, one by one.  After that, they are discovered correctly. 


One thing we learn, stay away from boot from SAN for ESX host, it takes 30 min for each host to come up after reboot even there are not a lot of luns in the ESX clusters.  It is proven much faster to boot ESX from local disk. 


Also, minimize the number of RDMs!!!  I work in another messy environment with lots of RDM.  Even rescan cannot happen during business hours.