With the exception of Role Based Access Control (RBAC), File Integrity Monitoring (FIM) is the only PCI requirement that achieves security in its purest form; prevention of, or alerts on, deviation from a known-good baseline.

Firewalls/Routers (DSS Req. 1.x) is almost there, configuration standards is even closer (DSS Req. 2.x), anti-virus (DSS Req. 5.x) is basically pointless, and logging (DSS Req. 10.x) could not be further off the mark. But FIM, assuming you have interpreted the requirements correctly, has real benefit when combined with the other requirements done well (especially configurations).

Unfortunately, even to this day, everyone associates Tripwire with the FIM requirement. If you can believe it, their NAME was included in version 1.0 of the PCI DSS! Of course, Tripwire licensing fees went through the roof as a result, and I’ve pretty much hated them for it ever since. Now they’re jumping on the Security Incident and Event Management (SIEM) bandwagon and doing as bad a job of it as everyone else.

PCI DSS v1.0 – “10.5.5 Use file integrity monitoring/change detection software (such a Tripwire) on logs to ensure that existing log data cannot be changed without generating alerts (although new data being added should not cause an alert).

Couldn’t even spell “as” correctly, but I digress.

As in all DSS requirements, the first question you must ask yourself is; “What is the intent of…?” In this case, the intent of FIM is to ensure all of your critical files (both operating system and application) do not change without authorisation (i.e. outside of a known change control scenario). Basically it’s seen by the SSC as a back-up for anti-virus (malware being the primary cause of unauthorised file changes), but in my opinion, the correct implementation of FIM, configuration standards, and baselined logging more of less negates AV altogether (see Annual Validation is Dead, it’s Time for Continuous ComplianceContinuous Compliance Validation: Why The PCI DSS Will Always Fall Short, and PCI – Going Beyond the Standard: Part 10, Anti-Virus for a little more background).

PCI DSS Requirement 11.5 states; “11.5 Deploy a change-detection mechanism (for example, file-integrity monitoring tools) to alert personnel to unauthorised modification of critical system files, configuration files, or content files; and configure the software to perform critical file comparisons at least weekly.” so it should be clear already how to go above and beyond; perform the checks more frequently that weekly. For a start, weekly is ridiculous, a lot can happen in a week, but the requirements’ very limited benefit is compounded by the fact that there is no guidance in WHAT changes you should be looking for.

FIMs can usually detect everything from file existence, size, permissions, hash values and so on, but these are only making checks against themselves from a previous ‘run’. Therefore one of the best ways to go WAY above PCI minimums is to compare the files to central database of known good configs directly from the operating system vendor themselves. Microsoft has a database of the latest and greatest system files (DLLS, EXEs etc.) against which you can run comparisons, and it should be relatively simple to add the baselines from each application you install on top.

Now let’s get REALLY crazy; What if you could then compare what files SHOULD be there as a result of a comprehensive configuration standard / hardening guide, and ensure that everything is as it should be?  Against CIS Security Benchmarks for example? Some FIMs (or similar agents) can also check Windows registry and GPO settings, so you can not only make sure everything is configured correctly per your approved standard(s), you can also automatically report against a significant number of other validation requirements (password complexity, access groups, log settings, time synch settings and so on).

And of course, the best way to blow PCI minimums out of the water is to compare a system against baselines stored centrally against each asset. These are system x’s available services, listening ports, permitted established connections and so on. Now you not only have configuration management, you have both policy and compliance validation built in automatically. Not once a year, but all day every day.

OK, I went WAY too far there, but hopefully you get the point. FIM is not just something you throw on a system because PCI demands it, you do it because it’s integral to how you do real security properly.

Yes Tripwire can do more than PCI asks for, but now you have just another management station to configure, maintain and monitor. FIM done well must integrate with all your other systems to have the necessary context, and FIM is only relevant if you have configured your systems correctly in the first place.

Finally, FIM should NEVER be seen as a stand-alone, end-point product, it can and should be a lot more than that.

Far too often, security is seen as a project, especially if PCI compliance is the goal. The requirements for vulnerability scanning and penetration testing are therefore seen as just another tick-in-a-box and their significant benefits lost.

External vulnerability scanning is the only requirement which must be outsourced and run by an approved scanning vendor (ASV, list here), the other requirements; internal vulnerability scanning, external penetration testing and internal penetration testing can by run by internal resources IF, and ONLY if, you can adequately demonstrate the requisite skill-sets in-house.

Of course, in order to save money, it is very tempting to skate by on the bare minimum, and unfortunately some security vendors (including QSAs) will allow you to do just that. Which is a shame, almost to the point of being irresponsible, as no other requirements give you a truer indication of your actual security posture than these.

Think of it this way; the bad guys use the EXACT same techniques to break into your systems that the good guys use to tell you what’s wrong. The ONLY differences between a hacker and an ethical hacker are intent and moral code, the skill-sets and mind-sets are the same.

Between vulnerability scanning and penetration testing, you have roughly 50% of your vulnerability management program sown up. Patch management management, risk management etc. make up the rest. However, the trick that’s almost always done poorly – if at all – is the integration of vulnerability management with asset management and change control. Any change to your environment should have appropriate vulnerability management processes around them, from a quick directed scan to a full blown credentialed penetration test, and all should be in-line with agreed configuration standards (as defined against each asset).

Going above and beyond PCI in scanning and pen. testing is relatively simple, but it’s not cheap in terms of resource cost. It also demands a maturity of process and a significant shift in culture to accept the ‘overhead’, but it’s more than worth it:

1. External Vulnerability Scanning – No choice but to use an ASV, but you should choose a vendor that provides 2 things at either no, or little, extra cost; Monthly scans (PCI requires quarterly), and unlimited directed scans (against single IPs, or subnets). Performed correctly, monthly scans and directed scans initiated by change control processes go significantly above and beyond. Note: For PCI do NOT open your external firewall/routing devices to your ASV’s IP addresses. Why would you decrease your security posture to test your security posture? Just run one scan for PCI, and THEN open your firewalls so that scanners can do a more thorough job. Keep these profiles separate, one for PCI only, one for your entire business.

2. Internal Vulnerability Scanning – You can do this yourself, and I’ve lost count of the number of clients running basic installations of Nessus, but unless you have significant expertise in how to configure it AND understand the results, don’t do it. For a start, any good QSA will fail you for lack of expertise, but do you really have the time to keep it up to date? Again, running internal scans monthly and as directed by change control goes above and beyond. Having two scan profiles is also a nice feature, but if the scan engine is capable of doing more than just rattle the windows (in the ubiquitous house analogy) and can actually perform a deeper scan / reconnaissance, then you have knocked this one out the park. PCI compliance is never security, do internal scanning as far above PCI minimums as you can afford.

3. External Penetration Testing – PCI requires that you attempt to break in (without breaking) via your Internet-facing presence, but poor guidance on what the test should consist of, combined with an enormous price-compression of pen. testing services means that this effort is usually more automated than I would consider appropriate. A pen. test is supposed to be a person with the necessary skills trying for days on end to discover ways into your systems. This is rarely the case now, but is EXACTLY what your should be doing. The Internet is where most breaches originate (used to be internal), so having a VERY robust security posture from the-outside-in is of paramount importance. Do NOT skimp on this one.

PCI calls for annual pen. tests, and to go above and beyond you need to perform these more frequently. This should not be an enormous cost, and most pen. test vendors can provide an infinitely scalable service based on scope and call-off days.

4. Internal Penetration Testing – Same premise as the external pen. test, but this time from the inside. PCI requires that this test simulate an attacker ‘plugging in’ where the admins sit and seeing what they can do from scratch. Above and beyond is therefore very simple; give the pen. tester FULL access to the environment, as well as credentials to go even further where appropriate. Like scanning, you have one test for PCI, then another test for your business.

There will be times when a simple vuln. scan of a system that has undergone change is not sufficient, so having a directed pen. test process available for critical business changes is very important.

None of the above processes should be stand-alone concepts, and should be very tightly integrated with risk assessment, change control and asset management processes to be truly effective. Vulnerability Management represents the end to each cycle of your security program (Plan > Do > Check > Act > Repeat), and ensures that your security posture always remain in-line with your business goals.

It bears repeating, do NOT skimp on this requirement, you will pay far more when you have to clean up the mess after a breach.

From everything I have seen in my many years performing PCI assessments, logging is not only one of the least understood of the requirements, it is the most under-utilised, and the one that gives/gave my clients the most pain.

Logging is the most important detective security control you have, bar none, and done correctly, logging is the foundation of your incident response program. Notice I didn’t say ‘disaster recovery’ as well, because if your incident response was where it should be you should not HAVE to recover from a disaster.

The confusion stems mostly from a lack of understanding of logging mechanisms themselves, even for Windows (for which PCI was clearly written). For example, do you think that Windows logs to the PCI requirements out of the box? I did too, but have been assured that it does not. Do you know HOW to get it to log appropriately? No, me either.

I have been further assured this if you WERE to turn logging on to cover the 10.2.X requirements, the logging would be so verbose as to render the device that’s doing the logging useless. Is that really the INTENT of logging? Of course not.

Also, can syslog EVER record the events required in 10.2.x, or even the event content as required in 10.3.x? Once again, I have been told no, but I am no expert.

Yes, you SHOULD have people who DO know this stuff, but how many organisations out there can truly afford that kind of deep expertise in-house? Yes you can outsource, but where is the guidance on EXACTLY how to configure operating systems to log to the PCI requirements? It probably exists, but in 10 years of doing PCI I have not found it, and I’ve even asked ‘experts’ in the field; Security Incident and Event Monitoring (SIEM) vendors. On that note, I have yet to see a SIEM vendor also be an expert in PCI, which to me is an absolute joke if that’s why they are selling it to their clients.

We have free hardening guides for Windows (CIS Security Benchmarks for example), but where is the guide that breaks down the Windows operating system into a mapping between the registry settings for logging and the PCI DSS? Or *nix flavours, or Cisco, or AS400, or Power Series? If you have them, please share?!

So let’s, for the sake of argument, assume that you cannot reasonably log to the letter of PCI, or more to the point, you do not WANT to for usability issues. Are you non-compliant? Let me answer that with another question; What’s more important, configuring your logging to record events for a forensics investigation, or configure your logging to help prevent a breach in the first place?

If you chose the second one, you are correct, and if you also choose to maximise your logging mechanisms, you will not only be PCI compliant to its intent, you will also be doing security properly.

First, logging is not about crunching masses of data through a correlation engine, its about the RIGHT data put into a base-lined context. There is no such thing as log event correlation without a deep understanding of what the end systems SHOULD look like, AND what your normal business processes are from start to finish. In other words, tell any vendor trying to sell you compliance though their SIEM, to put it where the sun don’t shine.

While we’re on the subject of SIEMs, how many of them do you think can accept Windows logs natively, or have to convert the logs to syslog via an agent (e.g. Snare)? Very few. What’s the point of buying a log mechanism that cannot even read WINDOWS events without butchering them DOWN to syslog?! You MUST ask the right questions before buying ANYTHING, especially a SIEM.

OK, so how DO you go above and beyond? Simple, in one way;

Do NOT perform your log reviews daily (10.6), because that’s just plain stupid, not to mention impossible to do adequately. Perform log reviews in real-time via some form of automation.

The automation you need is threefold;

  1. Events you should NEVER see: Each system admin, from OS, to network device to application SHOULD know which they events should never be seen under normal operating conditions. Look for these ‘strings’ and alert immediately.
    o
  2. Events you should not see in a certain quantity and velocity (i.e. thresholds): I don’t care if I see an admin fail to log in once, I do care if s/he fails 10 times in 2 seconds (for example).
    o
  3. Quantity of events over the course of time (i.e. trending): You have to save logs for 1 year (DSS 10.7), so why not put them to good use by trending events over time? Even if it’s just quantity of event (as opposed to quantity of type of event per device), the information you get can be extremely useful.

Perform all three of these things, and you have not only covered the ridiculous ‘daily reviews’ automatically, you now have input into your incident response mechanism that gives you real security.

Choosing the right centralised logging mechanism for your business is one of the most important decisions you can make, and it cannot be done ONLY for PCI. You must buy a system that can cover you enterprise-wide, and unless you have significant in-house expertise, you must build into the RFP the requirement for consulting support, and potentially some form of on-going managed service. Nothing stays the same, so your future state / needs will also need to be taken into account.

Do NOT penny-pinch here, but don’t buy anything that’s not appropriate. Your risk assessment process should tell you exactly what you need, and if you’ve not done one, start there.

The physical security requirements of the PCI DSS are by far the easiest to meet to the letter, but even these cause an inordinate amount of pain. This pain is rarely caused by the requirements themselves, but by the interpretation of them.

For example, if you were in-scope for the physical aspects, and I was to ask you if you HAD to have cameras to achieve PCI compliance, what would you say?

I you said something like “Not necessarily.”, “That depends.” or even a simple “No.”, you are correct. And if you don’t agree with that, just read the DSS;

“9.1.1 Use video cameras and/or access control mechanisms to monitor individual physical access to sensitive areas.

I’ve used the word ‘intent’ many times in my blogs about PCI, because the intent of a requirement is always more important than the words written. Unfortunately, a lot of my clients, or even their QSAs, don’t even read the words, let alone interpret the language into something the business side can understand.

The intent of the physical requirements is that you restrict access to only those who need it, and that you keep record of who was where, when. That’s it, and if THIS is too much, you have more problems than PCI.

Above and beyond is very simple, but there are too many options to go into here. Decide via your risk assessment process what is appropriate, get the relevant guidance if you don’t have it in-house and stick with that.

For back-ups, this is just as simple, and the requirements are written down for you. However, there are some gotchas that I will address here;

1. “9.6.1 Classify media so the sensitivity of the data can be determined.” does NOT mean that you have to LABEL the media AS sensitive, it just means that you have to have a way of distinguishing the media that may contain sensitive data so that you can protect it accordingly.

2. If you do have physical media that contains cardholder data, the decryption keys are not included, and the offsite facility has no means to GET access to the decrypt function, this media can be reasonably classified as out-of-scope for the same reasons P2PE is a de-scoping option for retail.

3.  Far more prevalent these days is some form of network access storage (NAS) where the data is ‘striped’ across many disks. From my perspective, because the theft of any one (or even several) disks does not constitute the theft of reconstitutable cardholder data, only the management station used to control the access to the data and NAS functions is relevant. Here all PCI controls relevant to a server would apply.

4. You must make sure you have a well defined Data Classification Policy or the data retention and destruction requirements have no context, and cannot be properly validated. You will probably end with a ton of data you don’t need and your overall risk will steadily increase over time.

5. Go back over all of your business processes that could have ever resulted in the retention of cardholder data and deal with the resulting media accordingly. This includes paper.

Finally, the new-to-v3.0 requirements for the protection of “devices that capture payment card data” (DSS Reqs. 9.9.X). Bottom line; if you don’t have this stuff in place already, I can’t help you, and you may want to consider alternate employment options.

Not much more to say here, but I don’t want to underplay this too much and make the mistake of the curse of knowledge.  knowing where your sensitive data is in ALL it’s forms is critical, and both your physical access to it, and the protection / destruction of it, are extremely important concepts.

There is no going above and beyond the PCI standard in this one, as every definition of Role Based Access Control (RBAC) is what is required. i.e. You must reduce the access to any system / location / data to only that required, based on the job responsibilities of the individual accessing it. Period / full stop.

So why doesn’t the PCI DSS just come out and SAY it must be RBAC? Because there are other ways of doing this that fall outside of the pure definition (a combination of Mandatory Access Control (MAC) and Discretionary Access Control (DAC) for example), and the SSC can never insist on something that may not be appropriate for all sizes of merchant and service provider alike. Just look at logging; you don’t HAVE to have a centralised logging mechanism (e.g. SIEM), but try performing the 10.5.X and 10.6 requirements without it.

The fact that there are not separate PCI standards for merchants and service providers, as well as separate standards for types of business (retail, e-comm, micro-merchant etc.) is why there is so much confusion, and why almost every self-assessment is BS. Yes, there are Self Assessment Questionnaires (SAQs), but these are reporting mechanisms only. Regardless of which SAQ you are asked to complete, you are still responsible for compliance to all DSS requirements. Combining all requirements, regardless of business, into one standard is like trying to describe an average human being. It’s the nuances that matter.

But I digress.

In practice, access control is almost universally performed badly, as it is seen as inefficient in terms of work product, too complicated to set-up, too labour intensive to maintain, or a combination of those three. Then there are the organisations who consider themselves too small to bother, or the ones for whom this is an alien concept. Yes, there are still some of those out there.

While you cannot go above and beyond this ultimate in security baselining, you can perform the function in a way that ensures that it automates much of the enforcement of RBAC, as well as produces the management information required to measure the appropriateness of the implementation.

Step 1 – Paperwork: Un-surprisingly enough, good access control starts with a Policy, is included in all relevant procedures, and is appropriately defined in relevant standards. Other than the Access Control policy itself, the most important paperwork foundation is the Data Classification Policy. RBAC should be tied to the level of the data involved, and without an understanding of data classification, this determination cannot be made. For example; a DBA who manages all financial or intellectual property, should have far more restricted access, and exponentially more applied oversight, than should the guy in charge of your public data.

Note: Steps 2 and 3 should be performed in parallel, and combined into an overarching management process that is managed though an Asset Management System (AMS);

Step 2 – Human Resources (or equiv.): Whatever unit in your organisation manages what has – until fairly recently anyway – been called HR, should have a complete understanding of the on-boarding procedures for every role within the organisation. While every employee begins with the exact same core procedures (paperwork, corporate training, assign email / domain access etc.), well defined roles will then continue along functional on-boarding road-maps. Depending on the organisation, there can be few or many of these, but either way, this needs to be generic enough to manage, but detailed enough to reduce the burden of requesting the remaining individual privileges required. For example; a Windows admin may be granted access to all Windows desktops on joining, but will have to earn the right to access critical Windows servers, and make a separate request through channels.

Step 3 – Asset Management: Asset Management is a critical aspect of access control. Well, asset management done correctly, gives you the ability to effect proper change control, because nowhere else in your organisation does the mapping of data / application / system to data classification occur to such an extent. Seeing as access control is enormously impacted by data classification, the only place where system / data / process ownership is fully defined is the asset register / database.

Step 4 – Enforcement: For every asset type, you need an enforcement mechanism around it. This will ideally be centralised (Active Directory, TACACS, RADIUS etc.), but sometimes local access lists are appropriate. 2 factor authentication (2FA) and/or separation of duties may be appropriate for critical systems, but not for guest wireless access, but whatever is chosen must be [yet again] appropriate in terms of sustainability, and security level.

Step 5 – Ongoing Management & Review: Over 90% of all processes fail because the necessary mechanisms were not put in place to sustain them. From management buy-in, to process definition, to initial implementation, to employee training, to culture shift, unless the process is part of an ongoing and enforceable life cycle it will die on the vine.

There are as many ways of effecting appropriate access control as there are organisations requiring it, so the above is as specific as I can get. How you implement this is your organisation WILL require a significant expertise, so if you don’t have that in-house, go find it.

If you don’t know what questions to ask, ask someone who does.