From everything I have seen in my many years performing PCI assessments, logging is not only one of the least understood of the requirements, it is the most under-utilised, and the one that gives/gave my clients the most pain.

Logging is the most important detective security control you have, bar none, and done correctly, logging is the foundation of your incident response program. Notice I didn’t say ‘disaster recovery’ as well, because if your incident response was where it should be you should not HAVE to recover from a disaster.

The confusion stems mostly from a lack of understanding of logging mechanisms themselves, even for Windows (for which PCI was clearly written). For example, do you think that Windows logs to the PCI requirements out of the box? I did too, but have been assured that it does not. Do you know HOW to get it to log appropriately? No, me either.

I have been further assured this if you WERE to turn logging on to cover the 10.2.X requirements, the logging would be so verbose as to render the device that’s doing the logging useless. Is that really the INTENT of logging? Of course not.

Also, can syslog EVER record the events required in 10.2.x, or even the event content as required in 10.3.x? Once again, I have been told no, but I am no expert.

Yes, you SHOULD have people who DO know this stuff, but how many organisations out there can truly afford that kind of deep expertise in-house? Yes you can outsource, but where is the guidance on EXACTLY how to configure operating systems to log to the PCI requirements? It probably exists, but in 10 years of doing PCI I have not found it, and I’ve even asked ‘experts’ in the field; Security Incident and Event Monitoring (SIEM) vendors. On that note, I have yet to see a SIEM vendor also be an expert in PCI, which to me is an absolute joke if that’s why they are selling it to their clients.

We have free hardening guides for Windows (CIS Security Benchmarks for example), but where is the guide that breaks down the Windows operating system into a mapping between the registry settings for logging and the PCI DSS? Or *nix flavours, or Cisco, or AS400, or Power Series? If you have them, please share?!

So let’s, for the sake of argument, assume that you cannot reasonably log to the letter of PCI, or more to the point, you do not WANT to for usability issues. Are you non-compliant? Let me answer that with another question; What’s more important, configuring your logging to record events for a forensics investigation, or configure your logging to help prevent a breach in the first place?

If you chose the second one, you are correct, and if you also choose to maximise your logging mechanisms, you will not only be PCI compliant to its intent, you will also be doing security properly.

First, logging is not about crunching masses of data through a correlation engine, its about the RIGHT data put into a base-lined context. There is no such thing as log event correlation without a deep understanding of what the end systems SHOULD look like, AND what your normal business processes are from start to finish. In other words, tell any vendor trying to sell you compliance though their SIEM, to put it where the sun don’t shine.

While we’re on the subject of SIEMs, how many of them do you think can accept Windows logs natively, or have to convert the logs to syslog via an agent (e.g. Snare)? Very few. What’s the point of buying a log mechanism that cannot even read WINDOWS events without butchering them DOWN to syslog?! You MUST ask the right questions before buying ANYTHING, especially a SIEM.

OK, so how DO you go above and beyond? Simple, in one way;

Do NOT perform your log reviews daily (10.6), because that’s just plain stupid, not to mention impossible to do adequately. Perform log reviews in real-time via some form of automation.

The automation you need is threefold;

  1. Events you should NEVER see: Each system admin, from OS, to network device to application SHOULD know which they events should never be seen under normal operating conditions. Look for these ‘strings’ and alert immediately.
    o
  2. Events you should not see in a certain quantity and velocity (i.e. thresholds): I don’t care if I see an admin fail to log in once, I do care if s/he fails 10 times in 2 seconds (for example).
    o
  3. Quantity of events over the course of time (i.e. trending): You have to save logs for 1 year (DSS 10.7), so why not put them to good use by trending events over time? Even if it’s just quantity of event (as opposed to quantity of type of event per device), the information you get can be extremely useful.

Perform all three of these things, and you have not only covered the ridiculous ‘daily reviews’ automatically, you now have input into your incident response mechanism that gives you real security.

Choosing the right centralised logging mechanism for your business is one of the most important decisions you can make, and it cannot be done ONLY for PCI. You must buy a system that can cover you enterprise-wide, and unless you have significant in-house expertise, you must build into the RFP the requirement for consulting support, and potentially some form of on-going managed service. Nothing stays the same, so your future state / needs will also need to be taken into account.

Do NOT penny-pinch here, but don’t buy anything that’s not appropriate. Your risk assessment process should tell you exactly what you need, and if you’ve not done one, start there.

The physical security requirements of the PCI DSS are by far the easiest to meet to the letter, but even these cause an inordinate amount of pain. This pain is rarely caused by the requirements themselves, but by the interpretation of them.

For example, if you were in-scope for the physical aspects, and I was to ask you if you HAD to have cameras to achieve PCI compliance, what would you say?

I you said something like “Not necessarily.”, “That depends.” or even a simple “No.”, you are correct. And if you don’t agree with that, just read the DSS;

“9.1.1 Use video cameras and/or access control mechanisms to monitor individual physical access to sensitive areas.

I’ve used the word ‘intent’ many times in my blogs about PCI, because the intent of a requirement is always more important than the words written. Unfortunately, a lot of my clients, or even their QSAs, don’t even read the words, let alone interpret the language into something the business side can understand.

The intent of the physical requirements is that you restrict access to only those who need it, and that you keep record of who was where, when. That’s it, and if THIS is too much, you have more problems than PCI.

Above and beyond is very simple, but there are too many options to go into here. Decide via your risk assessment process what is appropriate, get the relevant guidance if you don’t have it in-house and stick with that.

For back-ups, this is just as simple, and the requirements are written down for you. However, there are some gotchas that I will address here;

1. “9.6.1 Classify media so the sensitivity of the data can be determined.” does NOT mean that you have to LABEL the media AS sensitive, it just means that you have to have a way of distinguishing the media that may contain sensitive data so that you can protect it accordingly.

2. If you do have physical media that contains cardholder data, the decryption keys are not included, and the offsite facility has no means to GET access to the decrypt function, this media can be reasonably classified as out-of-scope for the same reasons P2PE is a de-scoping option for retail.

3.  Far more prevalent these days is some form of network access storage (NAS) where the data is ‘striped’ across many disks. From my perspective, because the theft of any one (or even several) disks does not constitute the theft of reconstitutable cardholder data, only the management station used to control the access to the data and NAS functions is relevant. Here all PCI controls relevant to a server would apply.

4. You must make sure you have a well defined Data Classification Policy or the data retention and destruction requirements have no context, and cannot be properly validated. You will probably end with a ton of data you don’t need and your overall risk will steadily increase over time.

5. Go back over all of your business processes that could have ever resulted in the retention of cardholder data and deal with the resulting media accordingly. This includes paper.

Finally, the new-to-v3.0 requirements for the protection of “devices that capture payment card data” (DSS Reqs. 9.9.X). Bottom line; if you don’t have this stuff in place already, I can’t help you, and you may want to consider alternate employment options.

Not much more to say here, but I don’t want to underplay this too much and make the mistake of the curse of knowledge.  knowing where your sensitive data is in ALL it’s forms is critical, and both your physical access to it, and the protection / destruction of it, are extremely important concepts.

There is no going above and beyond the PCI standard in this one, as every definition of Role Based Access Control (RBAC) is what is required. i.e. You must reduce the access to any system / location / data to only that required, based on the job responsibilities of the individual accessing it. Period / full stop.

So why doesn’t the PCI DSS just come out and SAY it must be RBAC? Because there are other ways of doing this that fall outside of the pure definition (a combination of Mandatory Access Control (MAC) and Discretionary Access Control (DAC) for example), and the SSC can never insist on something that may not be appropriate for all sizes of merchant and service provider alike. Just look at logging; you don’t HAVE to have a centralised logging mechanism (e.g. SIEM), but try performing the 10.5.X and 10.6 requirements without it.

The fact that there are not separate PCI standards for merchants and service providers, as well as separate standards for types of business (retail, e-comm, micro-merchant etc.) is why there is so much confusion, and why almost every self-assessment is BS. Yes, there are Self Assessment Questionnaires (SAQs), but these are reporting mechanisms only. Regardless of which SAQ you are asked to complete, you are still responsible for compliance to all DSS requirements. Combining all requirements, regardless of business, into one standard is like trying to describe an average human being. It’s the nuances that matter.

But I digress.

In practice, access control is almost universally performed badly, as it is seen as inefficient in terms of work product, too complicated to set-up, too labour intensive to maintain, or a combination of those three. Then there are the organisations who consider themselves too small to bother, or the ones for whom this is an alien concept. Yes, there are still some of those out there.

While you cannot go above and beyond this ultimate in security baselining, you can perform the function in a way that ensures that it automates much of the enforcement of RBAC, as well as produces the management information required to measure the appropriateness of the implementation.

Step 1 – Paperwork: Un-surprisingly enough, good access control starts with a Policy, is included in all relevant procedures, and is appropriately defined in relevant standards. Other than the Access Control policy itself, the most important paperwork foundation is the Data Classification Policy. RBAC should be tied to the level of the data involved, and without an understanding of data classification, this determination cannot be made. For example; a DBA who manages all financial or intellectual property, should have far more restricted access, and exponentially more applied oversight, than should the guy in charge of your public data.

Note: Steps 2 and 3 should be performed in parallel, and combined into an overarching management process that is managed though an Asset Management System (AMS);

Step 2 – Human Resources (or equiv.): Whatever unit in your organisation manages what has – until fairly recently anyway – been called HR, should have a complete understanding of the on-boarding procedures for every role within the organisation. While every employee begins with the exact same core procedures (paperwork, corporate training, assign email / domain access etc.), well defined roles will then continue along functional on-boarding road-maps. Depending on the organisation, there can be few or many of these, but either way, this needs to be generic enough to manage, but detailed enough to reduce the burden of requesting the remaining individual privileges required. For example; a Windows admin may be granted access to all Windows desktops on joining, but will have to earn the right to access critical Windows servers, and make a separate request through channels.

Step 3 – Asset Management: Asset Management is a critical aspect of access control. Well, asset management done correctly, gives you the ability to effect proper change control, because nowhere else in your organisation does the mapping of data / application / system to data classification occur to such an extent. Seeing as access control is enormously impacted by data classification, the only place where system / data / process ownership is fully defined is the asset register / database.

Step 4 – Enforcement: For every asset type, you need an enforcement mechanism around it. This will ideally be centralised (Active Directory, TACACS, RADIUS etc.), but sometimes local access lists are appropriate. 2 factor authentication (2FA) and/or separation of duties may be appropriate for critical systems, but not for guest wireless access, but whatever is chosen must be [yet again] appropriate in terms of sustainability, and security level.

Step 5 – Ongoing Management & Review: Over 90% of all processes fail because the necessary mechanisms were not put in place to sustain them. From management buy-in, to process definition, to initial implementation, to employee training, to culture shift, unless the process is part of an ongoing and enforceable life cycle it will die on the vine.

There are as many ways of effecting appropriate access control as there are organisations requiring it, so the above is as specific as I can get. How you implement this is your organisation WILL require a significant expertise, so if you don’t have that in-house, go find it.

If you don’t know what questions to ask, ask someone who does.

This may be the 12th post in my series, but I cannot stress the importance of good change control enough. I have said [too] many times that maintaining a good security posture is difficult enough, but to make things easier for bad guys from the INSIDE is just plain dumb.

If nothing in your environment changes, the only way risk can increase is by a change in the external threat landscape. Your Vulnerability Management processes should have this mostly covered.

Like almost every other process in security, change control should start with robust policy and procedure, and a well designed and up-to-date, Asset Management system. If you have not only a complete listing of your physical assets, but your business processes, data stores, system and data ownership, personnel & skill-sets, regulatory dependencies, and system criticality mapped out, you have the foundation for very effective change management. Difficult to create?; yes, difficult to maintain?; only if you don’t do it properly.

Speaking of change management, if you don’t have a change control board, get one. It does not matter what you call it, but it should contain enough of the right people to ensure that both the IT and the business side of the organisation can have full insight into the risk management process, and the upcoming changes to the environment, AND to make sure these changes are in-line with the business’s goals.

Good change control starts with a policy, is clearly explained in procedures, and is central to ANY changes that can affect the continued security of your information assets. From patching, to firewall rules, to application upgrades, to server on-boarding, everything must be reviewed and approved by the APPROPRIATE level of authorisation channel. Change control can never be seen as a bottle-neck or it will be bypassed, so unless you are able to rank your changes into distinct approval channels you will end up doing too much, or too little, neither of which is sustainable.

As far as PCI is concerned, you could email your change requests to whomever is responsible, and maintain the list on a spreadsheet. For small organisations this may be appropriate, but of all things to put online, standardise, and at least partial automate, change control is way up at the top of the list. Right next to your Asset Management System itself in fact. Or even better, PART of it!

Here’s how I think Change Control should be done;

  1. The ‘Paperwork’: You will need a –

 i.   Change Control Policy – Keep it simple, even something like; “All changes to IT systems (including, but not limited platforms, application, and data) will be subject to an appropriate review and approval process as determined by the highest data classification.”, is better than most organisations have in place.

ii.  Data Classification Policy – Without a data classification policy there is no way to determine the correct level of review and authorisation required. Patching will not go to your review board, but major application upgrades certainly will. Only appropriate processes are sustainable.

iii. Change Control Procedure – Specific and easy to follow instructions enabling EVERYONE in the organisation to request a change in a uniform manner.

  1. Change Control Board – It does not matter what you call this, Change Control Board (CCB), Change Advisory Board (CAB) or Rumpelstiltskin, the charter is the same; Bring to bear all relevant resources from the business and IT sides (and SMEs if appropriate) to review and approve all changes that meet the established criteria. This should include relevant system / data / process owners where appropriate.
    o
  2. Integrated Asset Management System (AMS) – Without an integrated AMS you lose the ability to add a second layer of criticality rating to your change requests (Data Classification being your first). A proper entry into the asset register will include process dependencies, regulatory implications, max. data classification, and a system ‘importance’ rank. When these assets are available to the change control requester as a drop-down list, the full impact of the requested change is immediately apparent, and the correct level of review and approval automatically assigned. If each asset also had the system / data / process owners assigned, the review board can be built and perhaps even alerted automatically.
    o
  3. Change Closure Checklist – Too often changes are closed before the correct testing is performed, but again, what testing and approval is appropriate should be tied to the importance of the change. Vulnerability scan results, penetration testing results, user testing input, and so on are just a sample of the things that could / should be done before any change is closed.
    o
  4. Periodic Review Against Management Metrics – The change was made for the benefit of the business in some fashion, right? How do you measure success? The answer of course is; “That depends.”, but a specific date / time should be set to review the changes made to ensure that all success criteria are met. This should not be complicated or labour intensive, ever, but the review process is the only way to feed back efficiency improvements into the risk management process. Nothing is perfect.

The above list assumes you have included the basic information required to request a change (Documentation of Impact, Back-Out Procedures etc.), but the PCI minimums are rarely sufficient for an organisation that really cares about this stuff. Again, you don’t want the request process to be ridiculously long or complex, but if you don’t provide enough information up-front, you cannot perform the correct review process, nor can you measure success of the results.

Finally, if your change control process is not owned by your Governance function, you will never be able to bring the correct business oversight to bear. IT and IT Security are only ever enablers, it’s the business side that needs to own the goals.

Another crazy long blog, apologies, but change control is just that important.

First, let me be clear; I hate anti-virus. I guess more accurately, I hate anti-virus companies who are still making squillions peddling their no-longer-relevant wares (in my opinion).

Blacklisting (i.e. signature based) end-point protection is meaningless and almost completely ineffective against zero-day attacks. It’s a game of constant catch-up that can (and will) never be won. Yet here we are, still buying anti-virus software because we don’t know better, and standards like the PCI DSS still call for it by name instead of dealing with the actual underlying issue.

What is the INTENT of anti-virus?  According to the DSS, you should;

5.1 Deploy anti-virus software on all systems commonly affected by malicious software (particularly personal computers and servers).

…and;

5.1.1 Ensure that anti-virus programs are capable of detecting, removing, and protecting against all known types of malicious software.

Commonly affected? As defined by whom? Clearly they mean Windows but can’t just come out and say it. Yes, other OSs are becoming increasingly affected by viruses, but would you call them common? More to the point; if you had to install and maintain anti-virus on all of your *nix and Apple products would YOU classify them as ‘commonly affected’?

No, neither would I.

The intent of anti-virus is sound; do not let bad stuff run on your systems. However, if you were doing security properly, would this not be basically redundant? Even PCI includes the means by which anti-virus becomes [in my view] excessive;

  1. Security Awareness Training (Req. 12.6) – If users were properly educated, a huge chunck of malware outbreaks would not happen in the first place. If your organisation does not have a very robust program for ongoing security training, they have missed the cheapest, and most effective security control that has, and will, ever exist. Ignorance is a choice, never an excuse.
    o
  2. Configuration Standards (Req. 2.x) – In my continuing theme of never backing up bold statements with actual facts, I will pronounce that the majority of malware out there is ONLY effective because the systems on which the malware is loaded are not configured correctly. Either the hardening guides are absent or inadequate, or the ongoing maintenance of the configurations was neglected.
    o
  3. Vulnerability Management (Req. 6.1) – If all you are relying on is patch releases from your OS vendors, then you deserve what you get. Vulnerability Management is everything from Patching, to Vulnerability Scanning, to Penetration Testing, to Change Control and Incident Response. Done well, vulnerability management is the only way you stand even half a chance of keeping up with the bad guys, but something I have personally never seen done well.
    o
  4. File Integrity Monitoring (Req 11.5) – Don’t buy Tripwire (I hate them too), but figure out a way to detect if a known-good file changes in some way. I have seen a client write an MD5 recursive hash on system32 and write the results to event logs for monitoring. It was free, and effective, but required significant expertise. All you’re trying to do here is make sure things stay the same, and it almost begs the question; Why have AV at all if you have FIM? This question becomes far more relevant the more of these points you master, but I will never negate the concept / cliché of defence-in-depth.
    o
  5. Logging & Monitoring (Req. 10.x) – In my opinion, nothing in your detective security portfolio is as important as this control, and can be used to create the most effective and ‘blanket’ compensating control for PCI there is. If you know what every system SHOULD be doing, anything NOT that is something to investigate. Daily review of log files is a farce, only real-time alerts triggered by base-line deviations makes sense, and should be the top of any organisation priorities to get right. Few do, and the majority of Managed / Cloud Security Services don’t do this either.
    o
  6. Incident Response (Req. 12.9) – Why bother being in business if you don’t intend staying in business? Incident Response can prevent an event from becoming a business crippling disaster, yet, like Vulnerability Management, is almost universally neglected. Do this one badly and I for one have no sympathy.

You should notice one unifying theme across all 6 of these controls; they have a significant process component, not technology. Most security is process, and yet PCI has driven more technology spend than all other compliance / regulatory standards in history combined (yes, that’s another fact-less statement, but I would be amazed if it wasn’t true). Anti-virus vendors, FIM vendors, logging vendors (and QSAs of course) have all made multi-millions from PCI, and not one of these vendors (including the QSAs) has ever made the effort to put their products into the proper context; A business focused solution that provides true benefit. Staying is business  IS an ROI!

All 6 factors will be addressed in their own Beyond The Standard posts, that should give some indication to their importance.

OK [deep breath], end of rant (and my longest blog of the series yet)! I’m not saying don’t use anti-virus if you believe it provides true benefit, and is not a massive capital / resource drain. But do NOT do it just because PCI says you should, do NOT rely on it, and focus your efforts on the above 6 factors as they are the things that actually meet the intent.

If you are only doing PCI minimums your QSA probably has no choice but to insist on AV (especially for Windows), your job is to give them an alternative.